What the first edition got wrong, and what changed.
The book's central instrument, the ratio R = L / H, was described with inconsistent arithmetic language, more than one green boundary, and an oversight half-life that did not mean what it needed to mean. This page records every correction, the one policy the book and the companion tools now share, and what is still to be done.
Publication note to existing readers
Two version numbers identify different things. The book corrections are issued as Edition 1 errata and method clarification 1.1; the book keeps "First edition" as its publication identity. The associated operational specification is 1.0.0. Page references below are printed page numbers of the 253-page first edition; add 14 for the PDF page number.
- Change register
- One ratio-band table (E1-02)
- Corrected arithmetic language (E1-01)
- Definitions and migration (E1-03)
- Qualification beside every result (E1-04)
- One veto and permit table (E1-05)
- The Commons claim (E1-06)
- How the framework should be judged (E1-07)
- RSI index consistency (E1-08)
- Factual and production corrections (E1-09, E1-10)
- What the site and tools now do
Change register
| ID | Type | Location | Action |
|---|---|---|---|
| E1-01 | Arithmetic correction | pp. 121-122, 209 and related prose | Correct numerator and denominator language against the formula actually in use |
| E1-02 | Policy reconciliation | pp. 119, 208-209 and site v0.1 | Retain the book's below-0.25 green threshold; retire the site's below-0.5 boundary |
| E1-03 | Substantive method clarification | pp. 119-123, 208-210 | Define remaining H and detection-inclusive L; require manual reassessment of legacy inputs |
| E1-04 | Claim qualification | Argument in Brief, Introduction, pp. 119-123, 188-190, 208-209, FAQ | Remove implications that a ratio proves control, safety or a timeless law |
| E1-05 | Policy revision | pp. 168, 205, 207, 226 | Replace conflicting veto summaries with one scope-specific decision table |
| E1-06 | Evidence qualification | pp. 151-152; note 59, p. 237 | Label the Commons claim by actual maturity and assurance evidence |
| E1-07 | Argument revision | p. 189, reply 6 | Replace opposing-camps validation with observable criteria |
| E1-08 | Internal consistency | p. 126 | Make the RSI table, prose and figure prescribe the same response |
| E1-09 | Factual correction | pp. 1, 174 | Separate Knight Capital's immediate financing crisis from its later corporate combination |
| E1-10 | Production and referencing | p. 127; p. 131; notes | Shorten the overflowing running head; correct a forward reference; improve source specificity |
E1-03 and E1-05 change operational interpretation. They are identified openly as method and policy revisions, not typographical corrections. The stricter band is retained for consistency and conservative handling, not because a validation study has shown 0.25 to be optimal.
E1-02One ratio-band table
The charts on pp. 119 and 208 and the band table on p. 209 are replaced by the following labels and boundaries. The same constants are used by the site, the assessment tool and every maintained integration.
| R used for policy | Band | Interpretation |
|---|---|---|
0 ≤ R < 0.25 | Green | Within the lower ratio band; all other permission checks still apply |
0.25 ≤ R < 1 | Amber | Reduced timing margin; explicit conditions and review required |
R ≥ 1 | Red | Outside the permitted timing margin for the requested scope |
| Not defensibly estimable | Unknown | Gather timing evidence; do not manufacture a value |
| Critical oversight evidence invalidated | Invalidated | The prior assessment does not support the current scope |
Where planning bounds are provided, the R used for policy is R upper = L upper / H lower, shown together with the central estimate and the range. Exactly 0.25 is amber, 0.40 is amber, 0.50 is amber, and 1.00 is red. Classify before rounding. "In control" and "control lost" are removed as categorical chart labels.
The history, for the record: Edition 1 and the first downloadable bundles used 0.25. Between June and July 2026 the companion site's hero demo and calculator used 0.7. The first published draft of the specification (v0.1, July 2026, archived unchanged) used 0.5. Any reading taken against 0.5 or 0.7 that landed between 0.25 and that line was called green then and is amber now. The upper boundary was never in dispute: R at or above 1 was red in every version.
E1-01Corrected arithmetic language
Where the formula is R = L / H, faster intervention reduces L, the numerator, and more durable oversight increases supported H, the denominator. Several passages had these the wrong way round or attached the labels to a different, conceptual change-rate relationship whose primitives differ. The book's "grow the denominator" and "shrink the numerator" language has been checked against the formula in use at each occurrence, and it is no longer claimed that every gear improvement changes either quantity numerically.
The hypothetical lab on p. 121 is restated: with an assumed H of four days and L of six days, R is 1.5. Its intervention time exceeds the stated oversight horizon, and the proposed activity requires restriction or redesign while the team investigates both quantities and the consequences that could occur before intervention. Where reviewers cannot establish the critical claims behind a requested authority, the oversight basis is no longer demonstrated; that gap is recorded explicitly, and a numerical ratio is not inferred from the loss of comprehension alone.
The worked calculation on p. 122, replaced
Suppose the evidence supporting an AI coding agent's narrowly specified activity is estimated to remain valid for three days, subject to immediate revalidation after listed changes. Detection takes half a day, the decision takes a day, and execution takes a day. Then L is 2.5 days and R = 2.5 / 3, approximately 0.83: amber.
Now give the accountable operator preauthorized intervention authority and test the execution path. Suppose detection still takes twelve hours, the decision takes five minutes, and execution takes twenty minutes. L is now twelve hours and twenty-five minutes. With H still assumed to be seventy-two hours, R is approximately 0.172: within the green policy band. This improvement comes from reducing the numerator. It does not by itself establish safety: the team must still examine what can become irreversible during those twelve hours, whether the critical evidence remains valid, and whether the other permission conditions are satisfied.
All numbers are hypothetical. The replacement also makes visible a lesson hidden in the original example: improving decision and execution time still leaves detection as the dominant delay.
E1-03Definitions and migration
The definitions on pp. 119 and 208 are replaced with the following. They are the same definitions the specification and the assessment tool use.
Intervention latency L is the time from the first point at which a specified deviation should be detectable under the documented monitoring design to confirmation that the required intervention has taken effect. It includes detection, decision, and execution delay. State the start and end events and identify the evidence supporting the estimate.
Oversight validity horizon H is the estimated remaining duration, at the time of assessment, for which critical evidence supporting the requested activity can reasonably be relied on before revalidation is required. Identify the critical claims, take the shortest defensible horizon among them, and list events that invalidate them immediately. If the horizon cannot be estimated, record unknown. If a critical claim is invalidated, the earlier assessment no longer applies.
The ratio R = L / H compares these quantities in the same units. It is one diagnostic input to a scoped decision. It does not measure the probability of harm, determine fairness, or establish safe operation.
Why H is no longer called a half-life
Earlier text described H as the time for half of a review's conclusions to become stale. That metaphor is replaced here with an explicit validity horizon because a single critical assumption may invalidate a decision before half of the reviewed material changes. This is a substantive clarification. Legacy half-life values must be reassessed rather than relabeled under the new definition.
Migration. An H recorded under the half-life definition, or under the general reliability wording of site v0.1, is not an H under 1.0.0. The tools do not convert it. A legacy assessment stays a read-only historical record with its original rule, raw inputs and classification until someone reassesses H and confirms the start event of L, and the drift tracker refuses to treat readings taken under different definitions as comparable. The workbook prompts now ask for the requested activity, the assessment time, the critical claims, lower, central and upper remaining H with its evidence and invalidation events, and the three L components with their start and end events. Seconds are recorded internally; readers enter familiar units.
E1-04Qualification beside every result
The opening promise of the measurement chapter, and the workbook's claim that the reader will know whether control remains, are replaced with:
This assessment helps expose mismatches between intervention time and the evidence supporting oversight. Read its timing result alongside critical hazards, the five gears, tested controls, and the authority to act. A favorable ratio supports further judgment; it cannot establish safe operation by itself.
In the objections and the FAQ, references to an invariant law or a proof of control become a proposed relationship to investigate, and the text states that usefulness and calibration remain open empirical questions. The false-precision discussion on pp. 171-172 is preserved, and a shorter qualification sits next to every standalone output, on the page and in the tools: "This assessment organizes evidence for a scoped decision. Its ratio band is a policy indicator, not a safety certification or permission to operate."
The green-ratio counterexample
If H is forty-eight hours and L is twenty minutes, R is approximately 0.0069. Yet an unacceptable disclosure might complete in one second after the same initiating event. A low ratio therefore does not replace prevention where consequences become irreversible before correction can take effect. The numbers in this example are illustrative.
E1-05One veto and permit table
The conflicting rules on pp. 168, 205, 207 and 226 are replaced by this single table. The unit of judgment is the requested activity and system boundary: a contained experiment, internal operation, limited or broad deployment, or an expansion of autonomy. It is published identically in section 5 of the specification and on the assessment tool.
| Precedence | Condition for the requested scope | Advisory result | Required action |
|---|---|---|---|
| 1 | Any critical claim invalidated; an unacceptable hazard has inadequate controls; or a required intervention test failed | Blocked | Withhold the requested permission. If already operating, the accountable owner applies the predefined containment or restriction plan. Consider harms from abrupt shutdown. |
| 2 | Any GEARS rating red, or known R upper at or above 1 | Blocked | Redesign, restrict or pause the requested activity and reassess. A differently bounded activity needs its own assessment. |
| 3 | No known blocker, but timing, a critical claim, a gear, stakes, current evidence, a required test, owner, scope or required independent review is missing | Insufficient evidence | Collect the named evidence. Do not grant new permission from this assessment. A separately authorized evidence-gathering experiment is possible. |
| 4 | Complete evidence, no blocker, and either amber timing or any yellow gear | Eligible with conditions | Specify each condition, owner, due time and verification test. An accountable person must record scoped permission before use. Missing conditions produce insufficient evidence. |
| 5 | Complete evidence, no blocker, green timing, and all five gears green | Eligible for scoped permission | An accountable person may grant the described scope and expiry. No automatic expansion of autonomy. |
Known blockers outrank missing information. All reds apply to the activity actually assessed. A veto on public deployment does not automatically prohibit a separately bounded research activity, which needs its own assessment; no gear is omitted because an activity is research. Human authorization is a separate record: awaiting decision, withheld, granted, granted with conditions, revoked or expired. A name typed into a client-side tool records an assertion, not a signature, a verified identity or an enforced permission.
A green ratio cannot compensate for a red gear. A serious known failure takes precedence over missing information. Existing operations require a planned response that also accounts for the harms of abrupt shutdown.
The summed-stakes threshold of 18 is removed from permission logic. Independent review is required when any of severity, irreversibility, scale, autonomy or power concentration is rated 4 or 5, when fundamental rights are affected, or when the system is RSI-relevant. These judgments route review; they do not numerically establish risk. Each dimension needs anchors and a rationale, with no default midpoint.
E1-06The Commons claim
The subheading on p. 151 becomes A proposed infrastructure for verifiable coordination. The readiness and reciprocal-reveal claims are replaced: the Verifiable Compute Commons illustrates a proposed infrastructure for making specified compute-related commitments more inspectable within a declared observation boundary. Whether a particular claim can be verified depends on the implementation, trusted components, coverage, threat model and the behavior of participants outside that boundary. A proposal for reciprocal disclosure does not by itself establish simultaneous disclosure, universal detection of non-compliance or effective enforcement. In the absence of implementation evidence, the discussion is to be read as a design proposal.
Every component claim on the companion page will carry two separate labels, inspected per claim and never treated as a single ladder:
| Dimension | Permitted labels | Evidence needed |
|---|---|---|
| Maturity | Proposal; prototype; pilot; operational within named scope | Design record; runnable version; bounded field use; supported ongoing use, respectively |
| Assurance | Not reviewed; self-tested; independently reviewed; independently reproduced | Named tests and artifacts; review scope and reviewer; or a reproduction record |
Unsupported component claims default to proposal / evidence not supplied for this claim. Note 59 on p. 237 is corrected to a specific dated source. The author's actual relationship to the project is to be confirmed and disclosed; no unverified declaration is inserted. Until the companion page carries these per-claim labels, read it as a design proposal.
E1-07How the framework should be judged
Reply 6 on p. 189 argued that criticism from opposing camps was evidence the framework was right. It is replaced. Criticism from opposing camps can reveal different weaknesses; it cannot establish that a framework is correct. Effective Equilibrium should be judged by what happens when people use it to make consequential decisions: whether the method exposes previously unexamined assumptions, gives reviewers enough evidence to challenge the decision, and leads to interventions that work under the conditions tested, recorded together with the time and cost of the assessment, the useful work delayed, and the burdens placed on people who did not choose the system. The framework would be weakened if independent users could not apply its definitions consistently, if favorable scores repeatedly concealed failed essential controls, or if it added substantial effort without improving decisions compared with existing practice. Its thresholds and procedures should then change. These are proposed evaluation criteria for the Edition 2 pilot plan, not results already obtained.
E1-08RSI index consistency
The table, prose and figure on p. 126 prescribed different responses at the same level. This mapping is used everywhere, including the digital companion. It describes proposed governance responses, not current legal licensing requirements, and it attaches no dates to level transitions.
| Level | Capability description | Framework response |
|---|---|---|
| 0 | No meaningful AI research contribution | Ordinary governance appropriate to the actual activity |
| 1 | Assistant-level research support | Defined human authority and ordinary evaluation of the actual use |
| 2 | Expert copilot across a research workflow | Heightened evaluation and disclosure of scope and limitations |
| 3 | Autonomous bounded AI research tasks | Independent review, restricted autonomy and monitored experiments |
| 4 | Substantial acceleration of successor development | Stronger containment, external review and coordinated oversight; consider formal authorization mechanisms |
| 5 | Improvement exceeds demonstrably reliable oversight | Activate the predefined restriction or pacing plan; no further autonomous expansion without a new basis for oversight |
The prose that grouped levels 0 to 2 under ordinary practice is deleted. An activity's ordinary hazards can require stronger controls than its RSI level suggests.
E1-09 and E1-10Factual and production corrections
Knight Capital (pp. 1, 174). The corrected text reads: "The trading failure occurred on 1 August 2012 and precipitated an immediate financing crisis. Knight Capital Group and GETCO later combined to form KCG Holdings on 1 July 2013." This separates the rescue financing from the later corporate combination. Any financing amount or precise rescue date is verified separately before it is included. Source: the SEC order, facts section.
Production. The running head on p. 127 is shortened to "The Recursive AI Playbook". The reference on p. 131 to the cyber-and-compute chapter is corrected to a chapter that follows. In the notes, broad references to reporting or histories are replaced with exact titles, publication dates, relevant sections or pages and stable links. A citation to a proposal cannot support an implementation claim.
What the companion site and tools now do
The specification is released as 1.0.0 with the definitions, the five bands, the decision table, the review routing and the version handling above; version 0.1 is archived unchanged. The assessment tool, the completed Harbor example on the homepage and the MCP server read one shared policy module and the same conformance fixtures, and each shows the specification version beside its result. A green band never fills in the human decision. Entries are processed in the browser and are not sent to our servers; reloading clears unsaved entries.
Not yet done: the manuscript source has to receive these edits and the PDF and ebook have to be regenerated, with reflowed tables, charts, page references and running heads inspected. The Commons page has to carry per-claim maturity and assurance labels. The full second edition remains in development. Unupdated integrations carry a visible compatibility notice rather than a claim of 1.0.0 support.
Where this is recorded
The threshold changelog is in section 4 of the specification. The definitions of L and H, with the migration note, are in section 2. The one veto and permit table is in section 5 and on the assessment tool. The machine-readable contracts are served beside the specification: the evidence pack schema, the Harbor example records, the conformance fixtures and the Python reference. The second edition will carry these as printed.