Working draft for community review

A shared instrument for staying in control.

The Open Control Surface Specification gives organizations a measurable, versioned language for tracking control as capability advances, so control drift becomes visible before it becomes an incident.

Version 0.1 Status draft Assessment time under 1 hour Licence CC BY 4.0
01

Purpose and scope

This specification defines a minimal, practical control surface for systems that can act, plan, or modify themselves at high velocity. It is designed for continuous use across successive model versions, so that control drift can be detected and corrected before R exceeds 1.

It is intentionally lightweight. Any organization should be able to run a full assessment in under an hour using the free Equilibrium Toolkit or equivalent internal tooling. The goal is not perfection on day one. The goal is a shared instrument that makes control drift visible and actionable across organizations and over time.

02

Core definitions

Three terms carry the whole specification. Use them exactly.

R = L / H
the control ratio

Decision latency L

Time from detection of anomalous or high-stakes behavior to effective intervention, whether that intervention is human or automated. Measured in consistent units: seconds, minutes, or hours.

Oversight half-life H

Time until the current validation, monitoring, or policy map of the system becomes untrustworthy, whether through capability change, distribution shift, or self-modification.

Control drift

Any sustained increase in R, or movement of any of the five gears toward red, between successive model versions or major capability jumps. Drift is a regression, not a nuance.

Effective Equilibrium

The state in which R remains below 1 while capability continues to increase, achieved by deliberately growing correction capability at least as fast as the system's action capability.

03

The four-step process

Mandatory on every major model progression. A full cycle is triggered by any major model release, any fine-tune that expands autonomy, and any addition of significant new tools or actions.

01

Measure the two clocks

How fast the system can act, and how fast you can understand and correct it.

02

Compute R

Decision latency over oversight half-life, in shared units.

03

Score the five gears

Governance, Equity, Aligned incentives, Resilience, Steering, each green, yellow, or red.

04

Re-validate the breakers

Re-test or re-commit every circuit breaker against the current version.

04

Measuring the control ratio

L is the observed or simulated time from anomaly signal to effective stop, throttle, or isolation. H is the estimated time until current oversight data or policies are no longer reliable, based on the rate of capability change or observed behavior drift. Both must be expressed in the same units, and the measurement method must be recorded and versioned with the model.

BandReadingWhat it obliges you to do
R < 0.5 Comfortable control margin Continue. Keep measuring on every progression.
0.5 ≤ R < 1 Acceptable, with active monitoring Proceed only with monitoring that is actually staffed and actually watched. Grow the denominator now.
R ≥ 1 Control lost The system is operating on an outdated map. Promotion and continued high-autonomy operation are vetoed until corrected.
On the boundary at 0.5. Below it, your understanding of the system outlives your reaction time by at least double, which leaves room for one correction to fail. Between 0.5 and 1 you retain control only as long as nothing surprises you, which is not a property you can rely on in a system whose capability is still rising.
05

Five gears scoring rubric

Each gear is scored green, yellow, or red, relative to the current model's capability surface. The gears are five levers on one number: they are how you grow the denominator.

GearCore questionGreenYellowRed
Governance Who can say stop, how fast, and on what authority? Clear, tested, sub-minute authority with automated enforcement possible Authority exists but is slow, fragmented, or frequently overridden No reliable or enforceable stop mechanism at the speed of the system
Equity Who bears the risk and who holds the upside? Material downside is aligned with decision rights and upside Partial alignment; some externalization of risk Upside privatized, downside largely externalized
Aligned incentives Does the reward for caution beat the reward for speed? Compensation, promotion, and metrics explicitly favor early detection and correction Mixed signals; speed still dominates in practice Caution is ignored or actively punished relative to velocity
Resilience What survives the failure you did not predict? Chaos-tested, graceful degradation, recovery paths for unknown unknowns Some fallbacks exist but known brittle points remain System is brittle to unpredicted failure modes
Steering Once launched, can you change direction or only watch? Verifiable post-deployment control: throttle, isolate, edit, reverse Possible but high-friction or human-only Observe-only mode once the system is acting

Veto rule

Any red gear on a high-stakes or self-improving system constitutes an automatic veto until remediated. Governance red combined with Resilience red is an especially strong veto signal.

06

Circuit breaker requirements

Circuit breakers are pre-committed, measurable if-then tripwires, written while the room is calm. Each one requires all five of the following.

  • An explicit measurable condition. A threshold someone could check without a debate about what it means.
  • A named owner. A person, not a team, not a rota, not a committee.
  • A defined deadline or automatic execution. The action fires on time or fires itself.
  • No-meeting authority. The owner can act without gathering further consensus.
  • Re-validation against the current model version. Tested on the version actually running.
Breakers that have not been re-tested on the current model version are considered non-existent. They do not count toward coverage, and they must not be cited in a disclosure. A brake designed for a system that no longer exists is not a brake.
07

Versioning and drift tracking

A single assessment is a snapshot, and snapshots hide trends. The point of this specification is the sequence. Every major model version must carry the following.

  • The previous version's R and gear scores.
  • The current version's R and gear scores.
  • The delta, in direction and magnitude, for both.
  • The list of circuit breakers that passed or failed re-validation.

Drift is a regression

A rising R, or any gear moving toward red between versions, is control drift and must be treated as a regression: triaged, owned, and fixed, not noted and shipped.

08

Recommended disclosure format

Voluntary public reporting. Organizations may publish a short, comparable control surface summary. Comparability is the point: six lines that mean the same thing at every company make control drift legible across the industry.

Format
Model: [name / version]
Date of assessment: [YYYY-MM-DD]
R: [value] ([In Control / Narrowing / Lost])
Gears: G:[color] E:[color] A:[color] R:[color] S:[color]
Breaker coverage: [X]% of critical paths with tested, owned breakers
Last full audit: [date]
Worked example
Model: atlas-4.2
Date of assessment: 2026-07-25
R: 0.71 (Narrowing)
Gears: G:green E:green A:yellow R:yellow S:green
Breaker coverage: 42% of critical paths with tested, owned breakers
Last full audit: 2026-06-30
Reading the gear line. The letters are field positions, not a ranking: G is Governance, E is Equity, A is Aligned incentives, R is Resilience, S is Steering. Note that R: on the third line is the control ratio, while the R: inside the gear line is Resilience. Parsers should treat the gear line as a single field.
09

Implementation notes

  • Reference implementation. The free Equilibrium Toolkit (scorecard, ratio calculator, drift tracker, circuit-breaker templates, disclosure generator) implements this specification and runs entirely in your browser. The MCP server exposes the same logic to any agent.
  • Internal tooling is fine. Organizations may use their own tools, provided the definitions, scoring criteria, and veto rules remain compatible with this specification.
  • This is a living document. Breaking changes increment the major version; clarifications increment the minor version. Implementations should state the specification version they target.
  • Free to adopt. Published under CC BY 4.0. Fork it, translate it, embed it in your own process. The only request is that changed versions do not keep the same version number.
10

How a specification like this becomes the default

Awareness does not change behavior; incentives do. Organizations will adopt continuous control-drift tracking when the cost of not doing it exceeds the cost of doing it. These are the leverage points that change that arithmetic, roughly in order of force.

i.

A standard measurement language

Once investors, insurers, customers, and auditors all use R and the five gears, internal teams can no longer hide behind a proprietary "safety review" that means something different at every company.

ii.

Public, comparable disclosure

A six-line summary published on a regular cadence. Companies that publish gain credibility; those that decline become the visible outliers. Comparability does the work that exhortation cannot.

iii.

Independent audit

Technical auditors working from the same open rubric turn a self-report into a verified claim, the way SOC 2 or ISO certification does. Rigor becomes a market signal rather than a cost centre.

iv.

Insurance and liability pricing

If a higher R or a red gear correlates with higher premiums or coverage exclusions, control work acquires a budget and an executive sponsor overnight. This is the strongest economic lever available.

v.

Diligence and procurement

One short clause in enterprise procurement and investor diligence: provide the current R, gear scores, and breaker coverage for the models we will depend on. A few large buyers asking is enough to move the rest.

vi.

A public incident taxonomy

Real control failures, mapped back to a specific R and specific red gears. This is what makes the cost of drift concrete instead of theoretical. The first entry is written up here.

Design principle

This succeeds when internal engineering and risk teams want the scorecard, because it gives them leverage against product pressure. Not because anyone outside is lecturing them.

What this must not become

A lobbying group

Pure advocacy invites politics and dilutes technical credibility. The instrument has to survive changes of government.

An unrunnable rubric

Any scoring system nobody can complete in under an hour will not be completed. Precision that costs adoption is not precision.

A closed certification

The standard stays open and the measurement method stays reproducible. A certification monopoly would recreate the problem it was built to solve.