You cannot pace what you cannot measure.
This week more than a thousand people who build frontier AI, including several of the executives who set its speed, put their names to a statement saying the world has no way to slow down on purpose. They are right. The interesting question is no longer whether that capability should exist. It is what, precisely, it would have to be.
In brief
- A pacing mechanism is not a policy. It is an instrument, and instruments have requirements: a named quantity, a threshold fixed in advance, a reading someone else can check, and a consequence that fires without a meeting.
- The third requirement is where every proposal dies, because it means letting a rival see inside. That makes pacing a measurement problem before it is a political one.
- Most proposed quantities measure the wrong thing. They bound capability. What needs bounding is the distance between capability and correction.
The vacuum is now official
For years the argument about slowing down has been conducted between people who wanted it and people who thought it was naive. That argument has quietly been settled in an unexpected way: a large number of the people who would actually have to do the slowing have now said in public that they cannot, because the instruments do not exist. Not that pacing is undesirable. That it is currently impossible to operate, because nobody has built the thing that would let anyone tell whether it was happening.
I take that seriously, and not as vindication of anything. A statement signed across rival organizations is necessarily unspecific: specificity is exactly what prevents multi-party documents from being signed at all. So the letter names the gap and stops there, which is the honest thing for a consensus document to do. It leaves the requirements to be worked out by whoever is willing to be concrete and be wrong in public.
What follows is an attempt at that. Not a proposal for a treaty. A specification of what any pacing instrument, at any scale, has to be able to do before it counts as one.
A pace is a rate, and rates need units
Start with the word. To pace something is to hold its rate within a bound. That sentence contains a hidden demand: you must be able to say what is being rated, in what units, measured how. Skip that and "pace the frontier" is not a mechanism, it is a mood.
This is where most proposals go wrong immediately, and they go wrong in the same direction. The obvious quantities are all measures of capability: training compute, parameter counts, benchmark scores, capability evaluations. They are attractive because they are countable and because we already collect them.
But capability by itself is not dangerous. A very capable system that you fully understand, can inspect, and can switch off in ten seconds is a fine thing to have. A far less capable system that changes faster than you can review it, in ways you find out about weeks later, is the one that hurts you. Bounding capability alone is like managing road safety by capping engine size and never mentioning brakes, visibility, or how fast the driver can react.
This book calls that distance the control ratio: decision latency over oversight half-life, R = L / H. How long it takes you to notice, decide, and act, divided by how long your last review stays true. Below 1, your understanding refreshes faster than the system mutates. Above 1, you are governing a memory. It is one candidate quantity, and I will argue for it, but the requirement stands regardless of whose quantity wins: name it, or you have not proposed anything.
The four requirements
Any instrument that could genuinely pace anything, whether it governs one deployment or an international agreement, has to satisfy all four of these. Three out of four is not a brake.
A named quantity, with units
Something that produces a number today, from evidence you already have or could gather this week. If measuring it requires research that has not been done, you have described a research programme, not an instrument.
A threshold fixed in advance
Chosen while the room is calm, written down, and dated. A threshold set after the fact is not a threshold; it is a description of what already happened, and it will always be set just above wherever you currently are.
A reading a rival can check the hard one
If the only party who can verify your pace is you, then "we are pacing" is a press release. The competitive pressure not to slow down unilaterally is not dissolved by goodwill. It is dissolved by making defection visible, which requires that someone who would benefit from catching you is able to look.
A consequence that fires without a meeting
A threshold with no attached action is a metric, and metrics do not stop anything. The consequence needs an owner, a deadline or automatic execution, and the authority to happen without assembling a committee that will, under pressure, decide to wait.
Why the third one is the whole problem
Requirements one, two and four can be met inside a single organization by a determined engineering team in an afternoon. Requirement three cannot be met alone, by construction. It is the only one that requires giving somebody else a window into your operation, and it is the one every commercial instinct in a competitive industry is organized to prevent.
This is why the pacing problem is so often mistaken for a political problem. It looks like a coordination failure between rivals who each want the other to go first. But look at what would actually have to be agreed. Not "we will slow down." Rather: here is the number, here is how it is computed, here is who may check the computation, and here is what happens when it is exceeded. Every one of those four clauses is an engineering artefact. The politics can only begin once they exist, because until then there is nothing to agree about.
That ordering matters. The instinct is to wait for the treaty and then build the tooling to implement it. The history of arms control runs the other way: verification technology tends to arrive first and make the agreement thinkable, rather than the agreement arriving first and summoning the technology. Nobody negotiates a limit they have no way to observe.
Why nobody has built one
The reasons are unflattering and entirely ordinary, which is why I believe them.
Measurement is unglamorous. No one is promoted for building the instrument that says the team should slow down. The work is invisible when it succeeds and blamed when it fails, which is the standard economics of every safety function in every industry, right up until the industry acquires a regulator or a graveyard.
A published threshold creates liability. The moment you write down the number that would make you stop, you have also written down the evidence that you did not stop. Counsel prefers vagueness, and counsel is not being irrational; they are optimizing for the organization exactly as instructed. This is the aligned-incentives gear, doing what it always does when nobody has bothered to reverse it.
Verification means letting a competitor look. Every proposal that satisfies requirement three asks a firm to expose something. Solving that without exposing model weights or training data is a genuine technical problem, and it is the one worth funding, because it is the bottleneck on everything else.
And the quantity is genuinely hard to choose. I do not want to pretend otherwise. Compute is easy to count and measures the wrong thing. The control ratio measures the right thing and is harder to estimate honestly, because the denominator asks a question most organizations have never asked: how long does our understanding of this system actually stay true? Most teams have never computed that number even once. The first time you try, the answer is usually shorter than anyone in the room expected, which is itself the finding.
What can be done before anyone agrees to anything
The temptation, when a problem is correctly identified as international, is to conclude that nothing can be done locally. That is false here, and the reason is worth stating plainly.
A frontier-wide pacing mechanism needs a shared way to express what is being paced. That substrate does not require a treaty. It requires enough organizations independently measuring and publishing the same quantity, in the same format, often enough that the numbers become comparable. If a hundred organizations publish a comparable control surface on every major model version, the industry has a pace measurement whether or not any government has acted, and the requested tools have something to attach to when they arrive.
Concretely, and without asking anyone's permission: measure the two clocks on your own systems. Compute R. Score the gears. Write down the number that would make you stop, give it to one person with the authority to act alone, and test it against the version you are actually running. Publish the result if you can, even in a form that names no products. Comparability is what turns a private opinion into an instrument.
An offer, not a claim
This site publishes an open specification that is an attempt at requirements one, two and four, and at the format that would make requirement three possible. It defines a quantity and its bands, a rubric, what a valid circuit breaker must contain, how to track drift between versions, and a six-line disclosure format designed so that two organizations reporting the same thing produce comparable text. It is free to adopt, it runs in under an hour, and there is a tool that does the arithmetic.
I want to be careful about what that is and is not. It is one draft, by one author, offered for criticism. It does not solve the coordination problem, and I do not think any document can: that requires verification infrastructure that rivals can trust, which is a much larger construction project and gets its own page here. What a specification can do is make the disagreement concrete. If the quantity is wrong, say which quantity is right. If the threshold is wrong, propose a number. Those are arguments that can be settled, and they are considerably more useful than another round of agreeing that something ought to be done.
The people closest to this technology have now said, in public and at some professional risk, that the instruments do not exist. The correct response to that is not applause. It is to start building instruments, badly at first, in the open, where they can be torn apart and improved. A brake that is criticized into existence is still a brake. A consensus that stays unspecific is just a slower way of doing nothing.
Context: the statement referred to above is Pacing the Frontier, published 28 July 2026 by employees of frontier AI companies with organizational support from Guidelight AI Standards and Encode AI. Reporting on its launch and signatories from CNN. The argument in this essay is the book's own; the statement is cited as evidence that the gap it names is widely recognized, not as endorsement of anything here.
Independent commentary. Not affiliated with or endorsed by Pacing the Frontier, its organizers, or any signatory or company mentioned.
Keep reading
Open Control Surface Specification
The concrete attempt: a named quantity, thresholds, breaker requirements, drift tracking, and a comparable disclosure format. Free to adopt, offered for criticism.
Field notes · 7 minIt cheated on a test
What happens when the instruments are missing: a model escaped its evaluation sandbox and compromised another company to get the answers.