A shared instrument for staying in control.
The Open Control Surface Specification gives organizations a measurable, versioned language for tracking control as capability advances, so control drift becomes visible before it becomes an incident.
Purpose and scope
This specification defines a minimal, practical control surface for systems that can act, plan, or modify themselves at high velocity. It is designed for continuous use across successive model versions, so that control drift can be detected and corrected before R exceeds 1.
It is intentionally lightweight. Any organization should be able to run a full assessment in under an hour using the free Equilibrium Toolkit or equivalent internal tooling. The goal is not perfection on day one. The goal is a shared instrument that makes control drift visible and actionable across organizations and over time.
Core definitions
Three terms carry the whole specification. Use them exactly.
Decision latency L
Time from detection of anomalous or high-stakes behavior to effective intervention, whether that intervention is human or automated. Measured in consistent units: seconds, minutes, or hours.
Oversight half-life H
Time until the current validation, monitoring, or policy map of the system becomes untrustworthy, whether through capability change, distribution shift, or self-modification.
Control drift
Any sustained increase in R, or movement of any of the five gears toward red, between successive model versions or major capability jumps. Drift is a regression, not a nuance.
Effective Equilibrium
The state in which R remains below 1 while capability continues to increase, achieved by deliberately growing correction capability at least as fast as the system's action capability.
The four-step process
Mandatory on every major model progression. A full cycle is triggered by any major model release, any fine-tune that expands autonomy, and any addition of significant new tools or actions.
Measure the two clocks
How fast the system can act, and how fast you can understand and correct it.
Compute R
Decision latency over oversight half-life, in shared units.
Score the five gears
Governance, Equity, Aligned incentives, Resilience, Steering, each green, yellow, or red.
Re-validate the breakers
Re-test or re-commit every circuit breaker against the current version.
Measuring the control ratio
L is the observed or simulated time from anomaly signal to effective stop, throttle, or isolation. H is the estimated time until current oversight data or policies are no longer reliable, based on the rate of capability change or observed behavior drift. Both must be expressed in the same units, and the measurement method must be recorded and versioned with the model.
| Band | Reading | What it obliges you to do |
|---|---|---|
| R < 0.5 | Comfortable control margin | Continue. Keep measuring on every progression. |
| 0.5 ≤ R < 1 | Acceptable, with active monitoring | Proceed only with monitoring that is actually staffed and actually watched. Grow the denominator now. |
| R ≥ 1 | Control lost | The system is operating on an outdated map. Promotion and continued high-autonomy operation are vetoed until corrected. |
Five gears scoring rubric
Each gear is scored green, yellow, or red, relative to the current model's capability surface. The gears are five levers on one number: they are how you grow the denominator.
| Gear | Core question | Green | Yellow | Red |
|---|---|---|---|---|
| Governance | Who can say stop, how fast, and on what authority? | Clear, tested, sub-minute authority with automated enforcement possible | Authority exists but is slow, fragmented, or frequently overridden | No reliable or enforceable stop mechanism at the speed of the system |
| Equity | Who bears the risk and who holds the upside? | Material downside is aligned with decision rights and upside | Partial alignment; some externalization of risk | Upside privatized, downside largely externalized |
| Aligned incentives | Does the reward for caution beat the reward for speed? | Compensation, promotion, and metrics explicitly favor early detection and correction | Mixed signals; speed still dominates in practice | Caution is ignored or actively punished relative to velocity |
| Resilience | What survives the failure you did not predict? | Chaos-tested, graceful degradation, recovery paths for unknown unknowns | Some fallbacks exist but known brittle points remain | System is brittle to unpredicted failure modes |
| Steering | Once launched, can you change direction or only watch? | Verifiable post-deployment control: throttle, isolate, edit, reverse | Possible but high-friction or human-only | Observe-only mode once the system is acting |
Veto rule
Any red gear on a high-stakes or self-improving system constitutes an automatic veto until remediated. Governance red combined with Resilience red is an especially strong veto signal.
Circuit breaker requirements
Circuit breakers are pre-committed, measurable if-then tripwires, written while the room is calm. Each one requires all five of the following.
- An explicit measurable condition. A threshold someone could check without a debate about what it means.
- A named owner. A person, not a team, not a rota, not a committee.
- A defined deadline or automatic execution. The action fires on time or fires itself.
- No-meeting authority. The owner can act without gathering further consensus.
- Re-validation against the current model version. Tested on the version actually running.
Versioning and drift tracking
A single assessment is a snapshot, and snapshots hide trends. The point of this specification is the sequence. Every major model version must carry the following.
- The previous version's R and gear scores.
- The current version's R and gear scores.
- The delta, in direction and magnitude, for both.
- The list of circuit breakers that passed or failed re-validation.
Drift is a regression
A rising R, or any gear moving toward red between versions, is control drift and must be treated as a regression: triaged, owned, and fixed, not noted and shipped.
Recommended disclosure format
Voluntary public reporting. Organizations may publish a short, comparable control surface summary. Comparability is the point: six lines that mean the same thing at every company make control drift legible across the industry.
Model: [name / version] Date of assessment: [YYYY-MM-DD] R: [value] ([In Control / Narrowing / Lost]) Gears: G:[color] E:[color] A:[color] R:[color] S:[color] Breaker coverage: [X]% of critical paths with tested, owned breakers Last full audit: [date]
Model: atlas-4.2 Date of assessment: 2026-07-25 R: 0.71 (Narrowing) Gears: G:green E:green A:yellow R:yellow S:green Breaker coverage: 42% of critical paths with tested, owned breakers Last full audit: 2026-06-30
R: on the third line is the control ratio, while the R: inside the gear line is Resilience. Parsers should treat the gear line as a single field.Implementation notes
- Reference implementation. The free Equilibrium Toolkit (scorecard, ratio calculator, drift tracker, circuit-breaker templates, disclosure generator) implements this specification and runs entirely in your browser. The MCP server exposes the same logic to any agent.
- Internal tooling is fine. Organizations may use their own tools, provided the definitions, scoring criteria, and veto rules remain compatible with this specification.
- This is a living document. Breaking changes increment the major version; clarifications increment the minor version. Implementations should state the specification version they target.
- Free to adopt. Published under CC BY 4.0. Fork it, translate it, embed it in your own process. The only request is that changed versions do not keep the same version number.
How a specification like this becomes the default
Awareness does not change behavior; incentives do. Organizations will adopt continuous control-drift tracking when the cost of not doing it exceeds the cost of doing it. These are the leverage points that change that arithmetic, roughly in order of force.
A standard measurement language
Once investors, insurers, customers, and auditors all use R and the five gears, internal teams can no longer hide behind a proprietary "safety review" that means something different at every company.
Public, comparable disclosure
A six-line summary published on a regular cadence. Companies that publish gain credibility; those that decline become the visible outliers. Comparability does the work that exhortation cannot.
Independent audit
Technical auditors working from the same open rubric turn a self-report into a verified claim, the way SOC 2 or ISO certification does. Rigor becomes a market signal rather than a cost centre.
Insurance and liability pricing
If a higher R or a red gear correlates with higher premiums or coverage exclusions, control work acquires a budget and an executive sponsor overnight. This is the strongest economic lever available.
Diligence and procurement
One short clause in enterprise procurement and investor diligence: provide the current R, gear scores, and breaker coverage for the models we will depend on. A few large buyers asking is enough to move the rest.
A public incident taxonomy
Real control failures, mapped back to a specific R and specific red gears. This is what makes the cost of drift concrete instead of theoretical. The first entry is written up here.
Design principle
This succeeds when internal engineering and risk teams want the scorecard, because it gives them leverage against product pressure. Not because anyone outside is lecturing them.
What this must not become
A lobbying group
Pure advocacy invites politics and dilutes technical credibility. The instrument has to survive changes of government.
An unrunnable rubric
Any scoring system nobody can complete in under an hour will not be completed. Precision that costs adoption is not precision.
A closed certification
The standard stays open and the measurement method stays reproducible. A certification monopoly would recreate the problem it was built to solve.