Edition 1 errata and method clarification 1.1

What the first edition got wrong, and what changed.

The book's central instrument, the ratio R = L / H, was described with inconsistent arithmetic language, more than one green boundary, and an oversight half-life that did not mean what it needed to mean. This page records every correction, the one policy the book and the companion tools now share, and what is still to be done.

Publication note to existing readers

This update corrects inconsistent arithmetic language, aligns the book and companion tools on one ratio-band policy, and clarifies the limits of the assessment. Under the shared rule, values below 0.25 are green, values from 0.25 up to 1 are amber, and values of 1 or more are red. A ratio of 0.40 is therefore amber. These are provisional policy bands, not validated safety thresholds. The updated method also makes the scope of permission explicit and uses a defined oversight validity horizon in place of the earlier, ambiguous half-life wording. Historical assessments should retain their original definitions and be reassessed before comparison with the new method.

Two version numbers identify different things. The book corrections are issued as Edition 1 errata and method clarification 1.1; the book keeps "First edition" as its publication identity. The associated operational specification is 1.0.0. Page references below are printed page numbers of the 253-page first edition; add 14 for the PDF page number.

Status of the printed book. These corrections are prepared for the manuscript source. The published PDF and ebook have not been regenerated as of 19 September 2026, and the original files and release history are preserved. The companion site, the assessment tool and the shared policy module already apply the corrected method.

Change register

Change register: identifier, type, location in the first edition, and action
IDTypeLocationAction
E1-01Arithmetic correctionpp. 121-122, 209 and related proseCorrect numerator and denominator language against the formula actually in use
E1-02Policy reconciliationpp. 119, 208-209 and site v0.1Retain the book's below-0.25 green threshold; retire the site's below-0.5 boundary
E1-03Substantive method clarificationpp. 119-123, 208-210Define remaining H and detection-inclusive L; require manual reassessment of legacy inputs
E1-04Claim qualificationArgument in Brief, Introduction, pp. 119-123, 188-190, 208-209, FAQRemove implications that a ratio proves control, safety or a timeless law
E1-05Policy revisionpp. 168, 205, 207, 226Replace conflicting veto summaries with one scope-specific decision table
E1-06Evidence qualificationpp. 151-152; note 59, p. 237Label the Commons claim by actual maturity and assurance evidence
E1-07Argument revisionp. 189, reply 6Replace opposing-camps validation with observable criteria
E1-08Internal consistencyp. 126Make the RSI table, prose and figure prescribe the same response
E1-09Factual correctionpp. 1, 174Separate Knight Capital's immediate financing crisis from its later corporate combination
E1-10Production and referencingp. 127; p. 131; notesShorten the overflowing running head; correct a forward reference; improve source specificity

E1-03 and E1-05 change operational interpretation. They are identified openly as method and policy revisions, not typographical corrections. The stricter band is retained for consistency and conservative handling, not because a validation study has shown 0.25 to be optimal.

E1-02One ratio-band table

The charts on pp. 119 and 208 and the band table on p. 209 are replaced by the following labels and boundaries. The same constants are used by the site, the assessment tool and every maintained integration.

The one ratio-band table: the value of R used for policy, its band, and its interpretation
R used for policyBandInterpretation
0 ≤ R < 0.25GreenWithin the lower ratio band; all other permission checks still apply
0.25 ≤ R < 1AmberReduced timing margin; explicit conditions and review required
R ≥ 1RedOutside the permitted timing margin for the requested scope
Not defensibly estimableUnknownGather timing evidence; do not manufacture a value
Critical oversight evidence invalidatedInvalidatedThe prior assessment does not support the current scope

Where planning bounds are provided, the R used for policy is R upper = L upper / H lower, shown together with the central estimate and the range. Exactly 0.25 is amber, 0.40 is amber, 0.50 is amber, and 1.00 is red. Classify before rounding. "In control" and "control lost" are removed as categorical chart labels.

The history, for the record: Edition 1 and the first downloadable bundles used 0.25. Between June and July 2026 the companion site's hero demo and calculator used 0.7. The first published draft of the specification (v0.1, July 2026, archived unchanged) used 0.5. Any reading taken against 0.5 or 0.7 that landed between 0.25 and that line was called green then and is amber now. The upper boundary was never in dispute: R at or above 1 was red in every version.

E1-01Corrected arithmetic language

Where the formula is R = L / H, faster intervention reduces L, the numerator, and more durable oversight increases supported H, the denominator. Several passages had these the wrong way round or attached the labels to a different, conceptual change-rate relationship whose primitives differ. The book's "grow the denominator" and "shrink the numerator" language has been checked against the formula in use at each occurrence, and it is no longer claimed that every gear improvement changes either quantity numerically.

The hypothetical lab on p. 121 is restated: with an assumed H of four days and L of six days, R is 1.5. Its intervention time exceeds the stated oversight horizon, and the proposed activity requires restriction or redesign while the team investigates both quantities and the consequences that could occur before intervention. Where reviewers cannot establish the critical claims behind a requested authority, the oversight basis is no longer demonstrated; that gap is recorded explicitly, and a numerical ratio is not inferred from the loss of comprehension alone.

The worked calculation on p. 122, replaced

Suppose the evidence supporting an AI coding agent's narrowly specified activity is estimated to remain valid for three days, subject to immediate revalidation after listed changes. Detection takes half a day, the decision takes a day, and execution takes a day. Then L is 2.5 days and R = 2.5 / 3, approximately 0.83: amber.

Now give the accountable operator preauthorized intervention authority and test the execution path. Suppose detection still takes twelve hours, the decision takes five minutes, and execution takes twenty minutes. L is now twelve hours and twenty-five minutes. With H still assumed to be seventy-two hours, R is approximately 0.172: within the green policy band. This improvement comes from reducing the numerator. It does not by itself establish safety: the team must still examine what can become irreversible during those twelve hours, whether the critical evidence remains valid, and whether the other permission conditions are satisfied.

All numbers are hypothetical. The replacement also makes visible a lesson hidden in the original example: improving decision and execution time still leaves detection as the dominant delay.

E1-03Definitions and migration

The definitions on pp. 119 and 208 are replaced with the following. They are the same definitions the specification and the assessment tool use.

Intervention latency L is the time from the first point at which a specified deviation should be detectable under the documented monitoring design to confirmation that the required intervention has taken effect. It includes detection, decision, and execution delay. State the start and end events and identify the evidence supporting the estimate.

Oversight validity horizon H is the estimated remaining duration, at the time of assessment, for which critical evidence supporting the requested activity can reasonably be relied on before revalidation is required. Identify the critical claims, take the shortest defensible horizon among them, and list events that invalidate them immediately. If the horizon cannot be estimated, record unknown. If a critical claim is invalidated, the earlier assessment no longer applies.

The ratio R = L / H compares these quantities in the same units. It is one diagnostic input to a scoped decision. It does not measure the probability of harm, determine fairness, or establish safe operation.

Why H is no longer called a half-life

Earlier text described H as the time for half of a review's conclusions to become stale. That metaphor is replaced here with an explicit validity horizon because a single critical assumption may invalidate a decision before half of the reviewed material changes. This is a substantive clarification. Legacy half-life values must be reassessed rather than relabeled under the new definition.

Migration. An H recorded under the half-life definition, or under the general reliability wording of site v0.1, is not an H under 1.0.0. The tools do not convert it. A legacy assessment stays a read-only historical record with its original rule, raw inputs and classification until someone reassesses H and confirms the start event of L, and the drift tracker refuses to treat readings taken under different definitions as comparable. The workbook prompts now ask for the requested activity, the assessment time, the critical claims, lower, central and upper remaining H with its evidence and invalidation events, and the three L components with their start and end events. Seconds are recorded internally; readers enter familiar units.

E1-04Qualification beside every result

The opening promise of the measurement chapter, and the workbook's claim that the reader will know whether control remains, are replaced with:

This assessment helps expose mismatches between intervention time and the evidence supporting oversight. Read its timing result alongside critical hazards, the five gears, tested controls, and the authority to act. A favorable ratio supports further judgment; it cannot establish safe operation by itself.

In the objections and the FAQ, references to an invariant law or a proof of control become a proposed relationship to investigate, and the text states that usefulness and calibration remain open empirical questions. The false-precision discussion on pp. 171-172 is preserved, and a shorter qualification sits next to every standalone output, on the page and in the tools: "This assessment organizes evidence for a scoped decision. Its ratio band is a policy indicator, not a safety certification or permission to operate."

The green-ratio counterexample

If H is forty-eight hours and L is twenty minutes, R is approximately 0.0069. Yet an unacceptable disclosure might complete in one second after the same initiating event. A low ratio therefore does not replace prevention where consequences become irreversible before correction can take effect. The numbers in this example are illustrative.

E1-05One veto and permit table

The conflicting rules on pp. 168, 205, 207 and 226 are replaced by this single table. The unit of judgment is the requested activity and system boundary: a contained experiment, internal operation, limited or broad deployment, or an expansion of autonomy. It is published identically in section 5 of the specification and on the assessment tool.

One veto and permit table: precedence, condition, advisory result and required action
PrecedenceCondition for the requested scopeAdvisory resultRequired action
1Any critical claim invalidated; an unacceptable hazard has inadequate controls; or a required intervention test failedBlockedWithhold the requested permission. If already operating, the accountable owner applies the predefined containment or restriction plan. Consider harms from abrupt shutdown.
2Any GEARS rating red, or known R upper at or above 1BlockedRedesign, restrict or pause the requested activity and reassess. A differently bounded activity needs its own assessment.
3No known blocker, but timing, a critical claim, a gear, stakes, current evidence, a required test, owner, scope or required independent review is missingInsufficient evidenceCollect the named evidence. Do not grant new permission from this assessment. A separately authorized evidence-gathering experiment is possible.
4Complete evidence, no blocker, and either amber timing or any yellow gearEligible with conditionsSpecify each condition, owner, due time and verification test. An accountable person must record scoped permission before use. Missing conditions produce insufficient evidence.
5Complete evidence, no blocker, green timing, and all five gears greenEligible for scoped permissionAn accountable person may grant the described scope and expiry. No automatic expansion of autonomy.

Known blockers outrank missing information. All reds apply to the activity actually assessed. A veto on public deployment does not automatically prohibit a separately bounded research activity, which needs its own assessment; no gear is omitted because an activity is research. Human authorization is a separate record: awaiting decision, withheld, granted, granted with conditions, revoked or expired. A name typed into a client-side tool records an assertion, not a signature, a verified identity or an enforced permission.

A green ratio cannot compensate for a red gear. A serious known failure takes precedence over missing information. Existing operations require a planned response that also accounts for the harms of abrupt shutdown.

The summed-stakes threshold of 18 is removed from permission logic. Independent review is required when any of severity, irreversibility, scale, autonomy or power concentration is rated 4 or 5, when fundamental rights are affected, or when the system is RSI-relevant. These judgments route review; they do not numerically establish risk. Each dimension needs anchors and a rationale, with no default midpoint.

E1-06The Commons claim

The subheading on p. 151 becomes A proposed infrastructure for verifiable coordination. The readiness and reciprocal-reveal claims are replaced: the Verifiable Compute Commons illustrates a proposed infrastructure for making specified compute-related commitments more inspectable within a declared observation boundary. Whether a particular claim can be verified depends on the implementation, trusted components, coverage, threat model and the behavior of participants outside that boundary. A proposal for reciprocal disclosure does not by itself establish simultaneous disclosure, universal detection of non-compliance or effective enforcement. In the absence of implementation evidence, the discussion is to be read as a design proposal.

Every component claim on the companion page will carry two separate labels, inspected per claim and never treated as a single ladder:

Two labels for each Commons claim: maturity and assurance, with the evidence each needs
DimensionPermitted labelsEvidence needed
MaturityProposal; prototype; pilot; operational within named scopeDesign record; runnable version; bounded field use; supported ongoing use, respectively
AssuranceNot reviewed; self-tested; independently reviewed; independently reproducedNamed tests and artifacts; review scope and reviewer; or a reproduction record

Unsupported component claims default to proposal / evidence not supplied for this claim. Note 59 on p. 237 is corrected to a specific dated source. The author's actual relationship to the project is to be confirmed and disclosed; no unverified declaration is inserted. Until the companion page carries these per-claim labels, read it as a design proposal.

E1-07How the framework should be judged

Reply 6 on p. 189 argued that criticism from opposing camps was evidence the framework was right. It is replaced. Criticism from opposing camps can reveal different weaknesses; it cannot establish that a framework is correct. Effective Equilibrium should be judged by what happens when people use it to make consequential decisions: whether the method exposes previously unexamined assumptions, gives reviewers enough evidence to challenge the decision, and leads to interventions that work under the conditions tested, recorded together with the time and cost of the assessment, the useful work delayed, and the burdens placed on people who did not choose the system. The framework would be weakened if independent users could not apply its definitions consistently, if favorable scores repeatedly concealed failed essential controls, or if it added substantial effort without improving decisions compared with existing practice. Its thresholds and procedures should then change. These are proposed evaluation criteria for the Edition 2 pilot plan, not results already obtained.

E1-08RSI index consistency

The table, prose and figure on p. 126 prescribed different responses at the same level. This mapping is used everywhere, including the digital companion. It describes proposed governance responses, not current legal licensing requirements, and it attaches no dates to level transitions.

Recursive self-improvement index: level, capability description and framework response
LevelCapability descriptionFramework response
0No meaningful AI research contributionOrdinary governance appropriate to the actual activity
1Assistant-level research supportDefined human authority and ordinary evaluation of the actual use
2Expert copilot across a research workflowHeightened evaluation and disclosure of scope and limitations
3Autonomous bounded AI research tasksIndependent review, restricted autonomy and monitored experiments
4Substantial acceleration of successor developmentStronger containment, external review and coordinated oversight; consider formal authorization mechanisms
5Improvement exceeds demonstrably reliable oversightActivate the predefined restriction or pacing plan; no further autonomous expansion without a new basis for oversight

The prose that grouped levels 0 to 2 under ordinary practice is deleted. An activity's ordinary hazards can require stronger controls than its RSI level suggests.

E1-09 and E1-10Factual and production corrections

Knight Capital (pp. 1, 174). The corrected text reads: "The trading failure occurred on 1 August 2012 and precipitated an immediate financing crisis. Knight Capital Group and GETCO later combined to form KCG Holdings on 1 July 2013." This separates the rescue financing from the later corporate combination. Any financing amount or precise rescue date is verified separately before it is included. Source: the SEC order, facts section.

Production. The running head on p. 127 is shortened to "The Recursive AI Playbook". The reference on p. 131 to the cyber-and-compute chapter is corrected to a chapter that follows. In the notes, broad references to reporting or histories are replaced with exact titles, publication dates, relevant sections or pages and stable links. A citation to a proposal cannot support an implementation claim.

What the companion site and tools now do

The specification is released as 1.0.0 with the definitions, the five bands, the decision table, the review routing and the version handling above; version 0.1 is archived unchanged. The assessment tool, the completed Harbor example on the homepage and the MCP server read one shared policy module and the same conformance fixtures, and each shows the specification version beside its result. A green band never fills in the human decision. Entries are processed in the browser and are not sent to our servers; reloading clears unsaved entries.

Not yet done: the manuscript source has to receive these edits and the PDF and ebook have to be regenerated, with reflowed tables, charts, page references and running heads inspected. The Commons page has to carry per-claim maturity and assurance labels. The full second edition remains in development. Unupdated integrations carry a visible compatibility notice rather than a claim of 1.0.0 support.

Where this is recorded

The threshold changelog is in section 4 of the specification. The definitions of L and H, with the migration note, are in section 2. The one veto and permit table is in section 5 and on the assessment tool. The machine-readable contracts are served beside the specification: the evidence pack schema, the Harbor example records, the conformance fixtures and the Python reference. The second edition will carry these as printed.