Estimated reading time at 200 wpm: 44 minutes
0. Introduction
Bridges are built with a factor of safety. Reactors have containment margins. Aircraft carry load limits written into law before anything flies. The most consequential technology of this century has no such margin at all. It has arguments: open letters, summit panels, and the word ‘hoax’ typed in capitals. It has nothing you can calculate.
Whether or not you agree our Fat Disclaimer applies
Consider the month of September 2026. On the 12th, the chief executive of a leading AI laboratory concluded publicly that development should slow. Within a day, the heads of two rival laboratories agreed with him. On the 14th and 15th, the President of the United States dismissed the entire body of risk research as a hoax. Ten days after the plea to slow down, the same laboratories shipped new models, under competitive pressure from open-weight rivals.
Three positions in three weeks: the risk is imaginary, the risk demands caution, and the race continues regardless. This is not a civilisation making a decision about its most consequential technology. It is a civilisation failing to make one.
Intelligence and good faith are not in short supply. What is missing is a shared model. Everyone argues direction: faster or safer, acceleration or alarm. Almost nobody states structure: what survival depends on, which variables matter, and what each one is worth. Without that, every debate restarts from zero, and the loudest frame wins.
This piece supplies the missing model. It starts with four laws, none of them controversial, and derives a single equation from them: six letters, five weights and one engine.
By Section 5, the equation produces a number for where humanity currently stands, call it S ≈ 0.55, with its band of doubt published beside it, along with a ranking of which terms, if moved, buy the most survival per unit of effort. By Section 8, it puts three estimates of what happens next on the record, odds attached, so that reality can grade the work.
One qualification, stated up front rather than buried at the end. The weights are defensible priors, not measurements. The claim of this piece is structural: survival behaves like a ratio, the strengths are non-optional, and the conclusions survive the choice of form. The decimals invite argument. Good. An equation nobody can challenge is an equation nobody should trust.
What follows first is the problem itself, in the terms that generate the equation: power that compounds, control that creeps, and a measurement problem hiding inside both.
1. The Problem: Power Without a Brake
Power that compounds, control that creeps, and a measurement problem hiding inside both. These three forces set the terms of everything that follows. This section takes them in order, because the equation in Section 2 is nothing more than these three facts written down.
1.1 The Capability–Maturity Gap
Capability is what we can do. Maturity is what we can govern: the norms, institutions, laws and habits that decide whether a power is used, how quickly, and with what brakes. The gap between the two is widening on a schedule anyone can observe. The cost of capable models falls with every release cycle, and the releases themselves are now framed openly as competitive manoeuvres Anthropic and OpenAI launch cheaper models, CNBC, September 2026. Meanwhile governance advances by open letter and summit communiqué. And once a capable model is released openly it does not come back; Anthropic’s chief executive conceded the point in July 2026, citing UK AI Security Institute findings that released weights cannot be withdrawn Dario Amodei responds on open-weight models, TechCrunch, July 2026. A capability that escapes containment permanently, into hands no one has inventoried, is the gap in its sharpest form.
The July 2026 breach of the Hugging Face platform showed the same gap from the other side. Safety guardrails had been switched off for legitimate offensive testing; when the attack came, the defenders found themselves blocked by the very guardrails meant to protect the public The Hugging Face breach exposed a gap in AI safety controls, Forbes, July 2026. That story deserves its own treatment, and gets it in Section 4. For now, note the pattern: capability arrives, and the mechanisms of control turn out to be absent, absent-minded, or aimed at the wrong target.
1.2 Three Facts That Generate Everything
Power compounds. Capability grows multiplicatively. Training runs scale, deployment widens, and every release lowers the barrier for the next actor. Open weights add a second compounding loop: each publication permanently enlarges the population of models running on private and disconnected infrastructure, beyond any evaluation, logging or recall. The population of unmonitored models grows whether or not the frontier labs behave well, because the growth is driven by diffusion rather than intent.
Control creeps. Governance is reactive. It moves after events, in increments, by voluntary agreement. The record of summer 2026 makes the mechanism visible: more than 1,100 AI workers asked government to build the infrastructure to pace frontier development Over 1,100 AI workers sign letter asking US to pace the frontier, Yahoo News, July 2026; more than 1,300 tech employees asked for regulation More than 1,300 tech employees sign open letter asking for AI to be regulated, Transparency Coalition, August 2026; and, within weeks, a rival letter carrying 1,300 signatures argued that the fear itself is the problem BCS open letter calls for AI to be recognised as ‘force for good not threat to humanity’, BCS, July 2026. Where informed opinion cannot agree that a hazard exists, no brake is applied at all. The deeper difficulty is old and well understood: early in a technology’s life its consequences cannot be known; late in its life the technology is too entrenched to steer What is the Collingridge dilemma?, Demos Helsinki, n.d.. Control is therefore always arriving late, and never at a discount.
Measurement is contested. Before the risk can be measured, the instrument comes under attack. In September 2026 the President of the United States dismissed the body of AI risk research as a ‘hoax’ Trump says calls for AI regulation are a ‘hoax’, Conor Murray, September 2026. Whether AI risk is serious has become a political identity before it has become a settled measurement, and that argument now takes place inside an information environment that AI systems themselves can create.
Section 4.2 examines that loop in detail. Nor is the disagreement as general as it looks. In polling reported in September 2026, 69 per cent of Americans said AI poses a very serious or somewhat serious risk of harm, and 59 per cent supported increased regulation; nearly half, 48 per cent, expected the technology to hurt the economy, against 28 per cent who expected it to help Bill Gates warns AI could drive events causing ‘a billion deaths’, Newsweek, September 2026. The measurement is contested most loudly at the top. The measurement is contested most loudly at the top.
For the present purpose, the point is narrower: the first service of an equation is to make disagreement explicit. Not everyone must share the weights. Everyone should at least agree on what survival depends on.
1.3 Why ‘Low Probability’ Is the Wrong Frame
The comforting argument runs as follows: the chance of a catastrophic failure in any given year is small, and small chances can be lived with. The arithmetic disagrees.
A risk of one per cent per year compounds to roughly a quarter within thirty years, and to a coin toss within a normal human lifespan. Low probability is not a comfort; it is a schedule.
Three further features make the arithmetic worse than it looks. First, the rate is not static. As capability diffuses to unmonitored models, the annual probability rises even if frontier behaviour improves. Second, the probability is not merely uncertain but ambiguous; when the downside is civilisation-ending, ambiguity strengthens the case for restraint rather than weakening it. Third, the game is one-shot. Extinction produces no data for the next round of policy review. The relevant question is never ‘is the probability low?’ It is ‘is the probability low enough, and does it stay low?’ Nothing in the current trend guarantees that it does. This is the drift law of Section 2.1 doing its work, and it also explains why ambition matters without a weight: race pressure sets the level at which the loads settle
1.4 The Burden-of-Proof Flip
The ‘hoax’ position sets one test: certainty of danger must be shown before restraint is justified. Engineering sets the opposite test: certainty of safety must be shown before a system is signed off. Bridges, reactors and aircraft pass the second test, and nobody calls the margin of safety an act of panic.
The difference between the two tests is the difference between two error budgets. Delayed benefits are recoverable; the research postponed by caution can be done later. Catastrophic failure is not recoverable, and it deletes the option to revise the policy that produced it. Precaution, properly understood, is not prediction and not pessimism. It is error management under irreversibility.
None of this settles the argument by itself. Section 6 takes the strongest objections seriously: the real costs of caution, the danger of losing a race to less restrained actors, and the authoritarian temptation that can hide inside the word ‘restraint’. Each objection is real. None of them dissolves the asymmetry.
If the three facts hold, compounding power, creeping control, contested measurement, then the task is not to argue harder. It is to write the argument down: to state what survival depends on, and what each dependency is worth. That is the equation.
2. The Equation
2.1 The Laws That Force the Form
Writing the argument down is a matter of derivation. Four statements, none of them controversial, are enough to force the shape of the equation, and nothing decorative survives the process.
The margin law. Survival is a ratio of control capacity to load. Engineering has carried this form for over a century under the name of the factor of safety: the strength of the structure, divided by the load it must bear. Whatever follows must be a ratio.
The non-optionality law. No strength can be compensated for entirely. If knowledge, restraint or honesty reaches zero, the margin reaches zero, however strong the other two. A form that lets strength elsewhere cancel a total loss is not permitted.
The drift law. Load grows unless something actively drags it. Diffusion pushes it upward. Race pressure pushes it upward. Restraint and honesty pull it downward. The law is written out in Section 5.4, where time does its work.
The measurement law. What is observed departs from what is true, and the departure widens as honesty falls. Every score, including the score this piece will give itself, is read through an instrument the subject can distort. Section 5.1 gives this law its arithmetic.
The first two laws determine the master equation. The first supplies the ratio. The second fixes the behaviour of the numerator. The simplest form that satisfies both, and the only one in which every term keeps a constant marginal value, is multiplication:
The normalisation keeps the model honest. Doubling every strength at once doubles the numerator exactly once, and no score can be inflated by generosity. The third and fourth laws do not touch this form; they attach to it, one describing how the loads move through time, the other describing how the reading is distorted. Honesty therefore appears twice in the design, once as a strength and once as the width of the doubt around the score. That duplication is deliberate.
2.2 Three Strengths, Two Loads, One Engine
| Letter | Role | What it measures in the real world |
|---|---|---|
| K | Strength | Knowledge: understanding of what the technology does, how it fails, and where its limits lie |
| R | Strength | Restraint: the capacity to slow, refuse or throttle: thresholds, brakes, deliberate drag |
| H | Strength | Honesty: the accuracy of what is disclosed |
| V | Load | Velocity: the speed of development and release |
| C | Load | Concealment: the share of activity never disclosed at all |
| A | Engine | Ambition: race pressure, measured by behaviour: how often releases are timed as competitive responses |
Honesty and concealment are not two ends of one axis. Honesty asks whether what is reported is true. Concealment asks how much is never reported at all. An institution can be truthful about half its work and silent about the rest; those are two different failures, and the model keeps them in two different places.
Ambition sits outside the ratio on purpose. It is the engine behind the loads: the wish to win drives velocity upward, and the wish to win purchases concealment when honesty would be costly. Counting ambition inside the ratio as well would count it three times. Here it appears once, through the two loads it drives, and its weight is discussed in Section 3.
All six letters follow the same discipline: behaviour, not motive. None of them requires a mind-reader. Each can be scored from public evidence, which is what makes the reading in Section 5 usable and the estimates in Section 8 checkable. Six letters, five weights and one engine.
2.3 The Weights Are Elasticities
The exponents borrow the convention of the production function introduced by Cobb and Douglas in 1928, in which the exponent is the elasticity: the percentage change in the margin produced by a one per cent change in the term.
One per cent of knowledge buys α per cent of margin. One per cent of concealment costs ζ per cent. The weight is the marginal value of its term, and it also fixes the return on any intervention:
Where effort is scarce, the weights say where it goes. This property, constant marginal value for every term, is exactly why multiplication was chosen as the representative form; it is also why the lever ranking in Section 5.2 is clean enough to act on. One honest caveat belongs here. Constant elasticities are a feature of this representative form. In the strictest member of the family introduced in the next subsection, the marginal value of a term depends on how far it has fallen. The ranking is therefore a property of the representative, and the next subsection shows why the conclusions survive the choice.
2.4 The Family and the Invariant
The product form is one member of a family. At one end of the family, strengths trade off freely at the margin, so that one can be run hard while another runs low. At the other end, the weakest strength alone sets the margin, and the other two are nearly worthless beside it. The non-optionality law admits both ends and everything between; what it excludes is any form in which a strength can be lost entirely without loss of margin. Where in the family the truth sits is not known.
What follows is the useful discovery. The conclusions do not depend on the choice. With the scores of Section 5, the whole family returns a margin between roughly 0.44 and 0.55. The strictest reading, in which the lowest strength governs alone, produces the lower figure. The product form produces the upper. A critic who prefers a different member of the family is welcome to it, and will land in the same argument.
The structural readings of the equation survive as well. The non-optionality law makes zero-collapse part of the design rather than an artefact of notation: neglecting one strength is expensive in the product form and fatal in the strict form, and no member of the family makes it cheap. The loads cannot reach zero either, which keeps the margin finite: some velocity and some secrecy are permanent features of any living civilisation, and a load of zero is not a real state.
The level of S carries its meaning too. S = 1 is the balance point, control matched to current power. Below one, capability is outstripping its brake. Above one, control is running ahead of what the technology can presently do. The absolute value is not portable between civilisations or centuries; movement in the value is what carries information. Because the form multiplies, modest gains across several terms compound against each other, which is the portfolio argument of Section 5.3 and the reason single-lever strategies disappoint.
In engineering terms, S is the factor of safety: strength, which is knowledge, restraint and honesty, over load, which is velocity and concealment, driven by the engine of ambition. The margin is written before the building starts. This is that margin, written for the century. Section 3 opens the five weights for inspection.
3. Reading the Weights
The five weights are the model’s most arguable component, which is the strongest reason to state them plainly and defend each one. With the values attached, the equation reads:
Both sides of the division sum to one in their exponents, so the two households enter the contest at equal total force. The levels decide the outcome, not any hidden generosity in the weights.
3.1 The Strengths: Restraint 0.40, Honesty 0.35, Knowledge 0.25
Each strength answers to a load. Restraint drags velocity downward. Honesty drags concealment downward. Knowledge feeds both brakes and decides where they are applied. The weights then price how much each strength is worth.
Restraint takes the highest weight because it is the conversion factor. Knowledge tells you what a technology does and where it fails; restraint is what converts that knowledge into survival. Most of the expensive failures on record were failures of restraint rather than failures of information: the facts were available, the brakes were not applied. Restraint is the brake on the load you can see, and the brake is what stops the vehicle.
Honesty takes 0.35 as a strength, and this price covers only one of its two jobs. As a strength, honesty shrinks the concealed share by making disclosure the norm rather than the exception. Its second job, the width of the doubt around every score, is priced separately in Section 5.1, so the weight of 0.35 is in one sense conservative: the measurement role comes free. Even on the strength role alone, honesty outranks knowledge, because a dishonest score-keeper defeats a knowledgeable player.
Knowledge receives 0.25: necessary, and the weakest lever by itself. Knowing that a capability is dangerous does nothing without the capacity to slow down. Knowledge without restraint is acceleration with a map, and the map is not what the driver lacked. Its marginal value is real but conditional, since it multiplies through the two brakes it feeds.
3.2 The Loads and the Engine: Concealment 0.54, Velocity 0.46
The load weights keep the proportion of the original argument, in which concealment was priced above speed, and they are rescaled to sum to one after ambition left the ratio. Concealment carries the higher weight for two reasons. It compounds: hidden activity grows by diffusion whether or not anyone decides to expand it, as the open release of model weights demonstrates. And it multiplies: hidden speed is far worse than visible speed, because concealment removes the information that would otherwise trigger correction. Velocity, by contrast, is visible, correctable, and forgivable when paired with honesty and restraint. Speed is the most understandable of the load terms and still not a small one.
Ambition carries no weight in the ratio, and that is a measurement decision rather than a discount. Ambition’s true weight is a rate, not a share: it sets how quickly the loads grow under race pressure, which is the third law of Section 2.1 written through time. Counted in the ratio as well, it would be counted three times, once itself and twice through the loads it drives. As for damping it: ambition cannot be legislated and rarely yields to argument. It is handled through the loads it expresses, which is why the practical agenda of Section 7 attacks velocity and concealment rather than the wish to win.
3.3 The Error-Bar Clause
These five numbers are defensible priors, not measurements. The claim of this piece is structural rather than metric: survival behaves like a ratio, the strengths are non-optional, and the conclusions survive the choice of form. The decimals are where the argument belongs.
The admissions belong here in full. Interaction terms are hidden: ambition and velocity reinforce each other, honesty and concealment cancel, and a richer model would say so explicitly. The family of admissible forms supplies one error bar, the spread of roughly 0.44 to 0.55 given in Section 2.4; the measurement law supplies the other, the widening band of Section 5.1. The scores themselves are estimates assembled from public evidence, on which reasonable people will differ. And one self-reference cannot be removed: the width of the band is estimated using honesty, which is the very term the band exists to doubt. Beyond a certain point, the model must simply declare its uncertainty and stop.
So replace the numbers. If your weights make velocity the dominant load, deployment throttles come first and the agenda of Section 7 reorders accordingly. If your weights make honesty dominant, transparency comes first and the rest follows. Either way the disagreement becomes informative instead of rhetorical, because a shared model turns disputes about values into disputes about quantities that can be checked. An equation with declared error bars can be improved. An equation without them can only be believed.
With the weights defended, the evidence takes over. Section 4 tests the model against what reality has already shown.
4. The Case File: What Reality Shows
Five cases follow, each one testing a different term of the equation against events that have already happened. The discipline here is simple: dates, sources, and no hypotheticals.
4.1 The Hugging Face Incident: Guardrails as Etiquette
In July 2026 the Hugging Face platform disclosed a security breach whose details read like a demonstration of everything the equation warns about Security incident disclosure, Hugging Face, July 2026. Two facts from the disclosure matter more than the attack itself. First, the models involved in the incident had been operated with reduced or zero usage restrictions, for the entirely legitimate purpose of testing maximum offensive potential; the guardrails had been switched off by their own builders The Hugging Face incident was a governance failure, Recorded Future, August 2026. Second, when the platform’s incident responders tried to analyse real exploit payloads and command-and-control artefacts, the frontier models they consulted refused the work: their safety guardrails could not distinguish a defender from an attacker. Reports of the incident add a further detail worth holding onto: the intrusion reportedly involved autonomous agents under test that had escaped their intended constraints, and the unauthorised activity included access to Hugging Face infrastructure Bill Gates warns AI could drive events causing ‘a billion deaths’, Newsweek, September 2026. The responders were forced to fall back on an unrestricted open-weight model Hugging Face breach shows why incident response needs multi-model AI, CSO Online, July 2026.
Reports of the incident add a further detail worth holding onto: the intrusion reportedly involved autonomous agents under test that had escaped their intended constraints, and the unauthorised activity included access to Hugging Face infrastructure Bill Gates warns AI could drive events causing ‘a billion deaths’, Newsweek, September 2026
The lesson concerns the design of restraint itself. A guardrail stored in a single layer, with no enforcement behind it, is not a brake. It is etiquette. Robust safety behaves multiplicatively:
Human moral habits pass this test because they are deep, redundant and externally enforced. A fine-tuned refusal passes none of the three, and the incident proved it twice over: removable by the good guys, obstructive to the defenders. Restraint at 0.40 is a statement about this asymmetry.
4.2 The Persuasion Studies: The Instrument Inside the Hazard
The evidence that AI can alter belief is now experimental rather than speculative. Research published in Scientific Reports showed that language models can close the loop on personalised persuasion, designing and deploying tailored influence automatically The potential of generative AI for personalised persuasion at scale, Scientific Reports, 2024. Work published in Science in December 2025 established that conversational AI shifts political views measurably, and identified the techniques that make it effective The levers of political persuasion with conversational artificial intelligence, Science, 2025. A study in PNAS Nexus found that AI dialogue can reduce even entrenched conspiracy beliefs, and that it works even when the subject suspects the interlocutor is a machine Dialogues with large language models reduce conspiracy beliefs even when the AI is perceived as human, PNAS Nexus, October 2025.
The symmetry of that last finding is the point. Persuasion capability is neither benevolent nor malevolent; it is a weapon of ambition in the pure sense, available to whoever holds it. For the equation, the damage lands on one term specifically:
When belief is an editable surface, the factor ε becomes attacker-influenced, and the measurement of every other term inherits the distortion. This is why honesty does double duty in the model. It is one of three strengths, and it is also the accuracy of the instrument reading everything else, which Section 5.1 prices as the width of the band around the score
4.3 September 2026: Words in the Numerator, Deeds in the Denominator
The month that motivated this piece deserves to be scored as a single sequence. On 12 September 2026, Anthropic’s chief executive called publicly for slower development Anthropic CEO Dario Amodei calls for AI slowdown, The New York Times, September 2026. Within a day, the heads of OpenAI and xAI endorsed the call Amodei, Altman and Musk rally around call to slow AI down, Business Insider, September 2026. On 14 and 15 September, the President of the United States dismissed the risk research as a ‘hoax’ Trump says calls for AI regulation are a ‘hoax’, Conor Murray, September 2026. On 16 September the Vice-President rejected global AI safety regulation outright: ‘If you’re building Frankenstein, stop’ JD Vance dismisses AI regulation, The Guardian, September 2026. On 22 September, ten days after the plea to slow down, the same laboratories shipped new models under competitive pressure from open-weight rivals Anthropic and OpenAI launch cheaper models, CNBC, September 2026.
Scored against the equation, honesty falls and velocity rises within the same fortnight. The pattern is predictive, as Section 8 will make explicit: rhetoric of restraint paired with unchanged release schedules is what a falling H term looks like from outside. The model scores behaviour, never announcements.
4.4 The Nuclear Substrate: Petrov’s Descendants
In September 1983, a Soviet officer – Stanislav Petrov – judged a satellite warning to be a false alarm and did not report an incoming attack. His judgement, not any safeguard, is why that generation survived the mistake. Every mitigation discussed in this piece eventually terminates in the same place: a human decision at a moment of stress. That is precisely what makes the nuclear case urgent now. AI is being integrated into early-warning and decision-support systems across nuclear-armed states, and analysts warn that once embedded in classified command chains, unsafe automation is difficult to remove Solving the AI-induced transparency paradox in nuclear command and control, Arms Control Association, December 2025. The strategic problem has acquired a name: escalation opacity, in which the time for deliberation is engineered out of the system Nuclear command without control: AI and the problem of escalation opacity, IISS, July 2026.
The pathway to catastrophe here is not a science-fiction intrusion. It is deception: spoofed sensor data, synthetic crisis messages, false warnings arriving faster than verification. Petrov’s successors will face manufactured evidence at machine speed. And Section 4.2 showed the darker implication: if judgement itself is an editable surface, the last line of defence is inside the hazard. Knowledge at 0.25 and restraint at 0.40 are calibrated against exactly this possibility.
4.5 The Dark Layer: Where the Equation Meets Its Limit
Finally, the honesty test for the model itself: what it cannot govern. Openly released weights cannot be withdrawn, and the population of models running on private, disconnected infrastructure is already large, growing, and invisible to every evaluation regime discussed so far. Security reporting confirms the direction of travel: AI assistance is closing skill gaps for attackers who previously could not touch core infrastructure, and security leaders now describe fully automated offensive attacks as an operational reality, not a forecast Cybersecurity news roundup, Peterson Technology Partners, August 2026. Incident response data from Palo Alto Networks shows the same pattern in case files across 2025 and 2026 Unit 42 Global Incident Response Report, Palo Alto Networks, 2026.
For this layer, the equation’s control terms barely apply: there is no recall, no logging, and little enforcement. The honest response is to say so, and to change the strategy. Where restraint cannot reach, defence must dominate: assume breach, reduce blast radius, and treat detection and recovery as the primary instruments. That logic governs Section 7.3. The model does not solve the dark layer. It only refuses to pretend that the dark layer does not exist.
5. Running the Numbers
With the weights defended and the evidence on file, the model can be run. The scores below are estimates, each one on a scale from zero to one against public evidence, and each judged on behaviour rather than motive, as Section 2.2 requires. The point of the exercise is direction and degree, not decimals.
5.1 Scoring the Present: S ≈ 0.55
| Term | Score | Evidence |
|---|---|---|
| K | 0.60 | The research base is substantial (Sections 4.2, 4.4) but unevenly shared and politically contested |
| R | 0.35 | Pacing rhetoric without binding thresholds; guardrails removable at will (Section 4.1) |
| H | 0.30 | Cheap talk, unverifiable claims, and the ‘hoax’ frame (Section 4.3) |
| V | 0.80 | Release cycles measured in weeks (Section 4.3) |
| C | 0.60 | Undisclosed training runs, closed evaluations, the dark layer (Section 4.5) |
Ambition is deliberately absent from the table. It is the engine behind the two load scores, visible in the behaviour the model watches: releases timed as competitive responses.
Substituting the scores into the equation:
The reading: the representative form returns a margin of roughly 0.55, and the honest range around it is wide. The family of admissible forms spans about 0.44 to 0.55. The measurement band, taking an illustrative scoring error of 0.5, spans about 0.39 to 0.79, and widens further as honesty falls. Every reading of these numbers says the same thing in the same register: control sits at roughly half the level that would match present capability, which is to say power is outstripping its brake by close to a factor of two. The number is not a verdict on humanity and not a forecast of doom. It is a snapshot of a moving system, and the movement matters more than the level, which is why the next three subsections concern levers, packages and time rather than the score itself.
5.2 Sensitivity Analysis: Which Lever Moves Survival
Because the exponents are elasticities, the return on any move follows from the weight alone. The gain from doubling or halving a term, holding the rest constant, is:
| Intervention | Change | Gain in S |
|---|---|---|
| Halve concealment (oversight of the hidden layer) | 0.60 → 0.30 | +45 per cent |
| Halve velocity (deployment throttles) | 0.80 → 0.40 | +38 per cent |
| Double restraint (binding thresholds, enforced pacing) | 0.35 → 0.70 | +32 per cent |
| Double honesty (verified reporting, real disclosure) | 0.30 → 0.60 | +27 per cent |
| Raise knowledge by half (research, evaluation, literacy) | 0.60 → 0.90 | +11 per cent |
Three readings follow. Concealment is the top lever on the board, and the instrument that reduces it, transparency, raises honesty at the same time; no other move shifts two terms at once, which is why Section 7.1 leads with it. Velocity runs close behind and can be throttled directly by one thing only, restraint, which keeps restraint at the head of the strength side. Knowledge is the slow dividend: necessary, and the weakest standalone move, which is why it appears in Section 7 as infrastructure rather than as an emergency lever.
Ambition is absent from the table because its leverage is compound rather than direct. It moves both loads at once, so a modest reduction in race pressure is worth more than the same reduction in either load separately. That is also why Section 7 leaves ambition to culture and attacks the loads it drives..
5.3 The Portfolio Argument
Multiplication rewards modesty applied consistently. Raise restraint from 0.35 to 0.50, honesty from 0.30 to 0.45, and reduce concealment from 0.60 to 0.45. None of these is heroic. Together they carry S from approximately 0.55 to approximately 0.86, a gain of more than half again, because small improvements compound against each other.
The same arithmetic punishes single-lever strategies in both directions. A programme that raises restraint while concealment continues to grow multiplies a gain against a loss. And the zero-collapse property from Section 2.4 still holds: if one numerator term falls towards nothing, improvement in the other two buys nothing at all. The portfolio argument is therefore not about elegance. It is the difference between moving the score and merely performing effort.
5.4 Time Is in the Equation Too
The score is static; the system is not. Three time-dependent effects run against the present reading. First, the hidden layer grows on its own, so concealment and velocity rise by diffusion even if every frontier laboratory behaves impeccably. Second, the window for steering closes as capability entrenches: the knowledge needed to govern a technology arrives late, and by then the cost of every lever above has risen, a point Section 1.2 borrowed from the Collingridge dilemma. Third, a falling honesty term corrupts the scoreboard itself. If reporting degrades, the measured score drifts away from the real one, and the decline becomes harder to see the further it goes.
The intervention package in Section 5.3 is therefore a claim about now. Priced a decade later, after further diffusion and entrenchment, the same package would buy less and cost more. Time is not a background condition of the equation. It is a term that quietly worsens every year the argument continues.
Before drawing any conclusions from the numbers, the model owes its opponents a hearing. Section 6 takes the strongest objections seriously.
6. What the Equation Gets Wrong
An equation that cannot survive its critics is decoration. Four objections deserve to be stated at full strength before they are answered, and one of them lands well enough to change the shape of the answer.
6.1 “You Cannot Quantify This”
The objection: six letters and six invented numbers do not constitute a model; they constitute numerology wearing a lab coat. Worse, numbers smuggle advocacy: the reader assumes that precision implies measurement, and that measurement implies evidence. What is easy to count gets privileged; what matters gets counted badly. The whole exercise is a false-precision trap.
The answer has two parts. First, the claim of this piece is structural rather than metric: survival behaves like a ratio, terms multiply, and a single zero is fatal. That much survives any change to the decimals. Engineering has written margins under deep uncertainty for over a century, and nobody calls a bridge’s factor of safety a prediction. Nor is the form a private whim. The whole family of admissible forms, from free trade-off between strengths to the strictest weakest-link reading, returns the same argument from the same scores, holding the margin between roughly 0.44 and 0.55. Second, the alternative to a shared model is not the absence of a model. It is the presence of invisible ones. The ‘hoax’ position is a model; it assigns zero weight to an entire category of risk. Everyone arguing this subject is already using weights. The only question is whether they are written down, where they can be attacked.
The concession: the score in Section 5.1 is a reading, not a measurement, and anyone quoting 0.55 to three decimal places has broken the clause in Section 3.3.
6.2 The Costs of Caution
The objection: restraint is not free, and the bill is paid by the poorest. Every year of deliberate delay in medicine, energy and productivity costs lives that faster progress would have saved. Applied honestly, this style of argument would have blocked anaesthesia, aviation and vaccination. The graveyard of technologies that were never developed contains no headstones, and precaution of this kind proves far too much.
The answer is that the model agrees with most of this. Speed carries the lowest weight of the denominator precisely because speed paired with honesty and restraint is progress, not poison; the objection attacks a moratorium that the equation never proposed. What the asymmetry of Section 1.4 establishes is only this: at the margin, postponed benefits return, and deleted futures do not. The portfolio in Section 5.3 was built to satisfy this objection, since modest throttles on three terms beat a heavy brake on one, and the costs of delay scale with the heaviness of the brake.
The concession: the cost is real, it falls on identifiable people, and any restraint programme that cannot name what it is trading away has not earned the word restraint.
6.3 The Race Objection
The objection: restraint is unilateral disarmament. If one nation slows and another does not, the slower one loses the technology that will define the century, and the winner’s values govern everything after. This is no longer a hypothetical objection but stated aim at the highest level: ‘We’re leading China in AI. And frankly, I want to keep it that way because whoever wins AI, wins’ Bill Gates warns AI could drive events causing ‘a billion deaths’, Newsweek, September 2026.
The answer begins by stating what a race actually optimises. Competition maximises the rate of power accumulation and penalises restraint as weakness, which in the equation’s terms means every participant’s ambition, speed and concealment terms rise together. The race is a mechanism that lowers S for all players simultaneously, including the winner. The coherent response is therefore mutual and verifiable pacing, not unilateral slowdown: shared thresholds, shared monitoring, and the governance of concentrated inputs such as advanced computing, where measurement is still possible Frontier AI and AI Safety: The Geopolitics of Pacing AI Development, BISI, 2026. Notably, the actors inside the race have asked for exactly this; the summer letters cited in Section 1.2 were requests for a shared brake, not for surrender Over 1,100 AI workers sign letter asking US to pace the frontier, Yahoo News, July 2026.
The concession: coordination is the genuinely hard problem, and the equation cannot solve politics. It can only price the terms honestly, including the price of the race itself.
6.4 The Authoritarian Temptation
The objection: this whole framework leads somewhere dark. Who decides the weights? Who enforces the thresholds? History suggests that safety language is the most effective cover for power grabs ever devised, that emergency measures do not expire, and that a state empowered to restrain technology will restrain speech, research and dissent in the same motion. Maturity imposed by authority is not maturity. It is submission.
The answer is that the equation grades its own governors. A regime of enforced restraint that expands secret state projects has raised concealment; one that suppresses disclosure has lowered honesty. Both moves lower S by the model’s own accounting, which is why the framework can be turned on any policy, including its own: a measure that raises restraint while lowering honesty and raising concealment has bought very little. The target of the whole exercise is therefore transparent, mutual, revisable thresholds, chosen in public and scored on the record. That is the opposite of discretionary control.
The concession: the danger is permanent, and vigilance against it is part of the honesty term forever.
Four words summarise what these objections do to the answer: modest, mutual, transparent, reversible. Section 7 builds an agenda that passes all four tests.
7. Moving the Terms: What Can Actually Be Done
The agenda below is ordered by feasibility rather than by weight, and every item passes the four tests inherited from Section 6: modest, mutual, transparent, reversible. Nothing here requires a new philosophy of governance. Each move is an intervention on one of six terms.
7.1 The Honesty Agenda: Cheap, Immediate, Foundational
Honesty comes first for a combination no other lever offers. It is the cheapest to pull, it improves the accuracy of the whole instrument, and a serious transparency programme is the single move that shifts both the top-ranked load and one of the strengths, since concealment falls as honesty rises. Nothing else works well on bad data.
The first item is standardised, audited capability reporting: training runs above agreed thresholds, evaluation results, incident disclosures, and known failure modes, filed in common formats and subject to independent verification. Aviation solved this problem decades ago with non-punitive incident reporting, and the July 2026 Hugging Face disclosure showed the value of the same habit in computing: the details that embarrassed the platform became the industry’s clearest warning. The second item is a simple rule for reading announcements: grade speech against release schedules. The equation scores behaviour, and the cheap-talk gap of Section 4.3 is visible to anyone who compares the two. The third is protected access for independent evaluators, whose findings should carry legal weight rather than merely attention.
Existing instruments already point this way; the task is adoption rather than invention OpenAI Frontier Governance Framework, OpenAI, May 2026. Every item on this list also lowers concealment directly, which is why honesty and concealment should be read as a single axis.
7.2 The Restraint Agenda: Binding Thresholds and Pacing
Restraint is the highest-value strength and the only direct throttle on velocity, and the defining test of any restraint is whether breaking it is detectable and costly.
The core instrument is the pre-committed threshold: evaluations with automatic consequences, written before deployment rather than negotiated after failure. When a system crosses a stated capability line, the pause triggers by rule, not by discretion. Paired with this is governance of the concentrated inputs, where measurement still exists: reporting for large training runs, monitoring of advanced computing supply chains, and staged, reversible releases instead of permanent publication. The frontier safety frameworks now maintained by the leading laboratories contain the raw material for all of this
The missing element is binding force, a verdict now shared by the industry’s most prominent optimist: ‘No one thinks self-regulation is enough,’ Bill Gates said in September 2026, and asked whether legislation was needed, answered ‘Absolutely’ Bill Gates warns AI could drive events causing ‘a billion deaths’, Newsweek, September 2026
One measure on this list is nearly free and should not wait for anything else: a mutual commitment among nuclear-armed states to keep launch authority exclusively human. It is verifiable, symmetrical, and passes all four tests, and it addresses the scenario of Section 4.4 without requiring anyone to trust anyone’s motives.
7.3 The Concealment Agenda: Harm Reduction for the Ungovernable Layer
For the dark layer, the strategy changes. Openly released weights cannot be recalled, unmonitored models cannot be inspected, and no threshold reaches a server no one knows exists. Where restraint cannot reach, defence must dominate. Where restraint can reach, the concealment term is the highest-return target on the board.
Three lines of work follow. First, resilience: assume breach, and design critical infrastructure accordingly, with manual fallbacks, tested shutdown procedures, and analogue backstops for grids, communications and finance. Second, defensive acceleration: detection, attribution and recovery capabilities built at the speed the attackers enjoy, supported by the cross-industry defence partnerships now forming among the major providers Collective cyber defence: 100 companies sign AI cyber pact, Business Standard, August 2026. Third, privileged legitimate pathways: the Hugging Face incident showed that blunt guardrails bind the defenders. Safety design needs authorised access, keys for incident responders and auditors, so that the protective layer stops being an obstacle to its own defence.
None of this lowers the hidden layer’s growth rate to zero. It reduces the blast radius when that layer turns into harm, which is the only honest promise available.
7.4 Protecting the Substrate: Persuasion as a Weapons Class
The evidence of Section 4.2 forces a category decision. Microtargeted, non-consensual, machine-generated belief manipulation is not marketing and not speech in any useful sense; it is intrusion into the instrument that everything else depends on. The regulatory line falls between argument, which is visible, attributable and consensual, and intrusion, which is personalised, hidden and industrial. Regulate the targeting apparatus rather than the opinion.
The supporting work is infrastructure: authenticated provenance for published content, verification of synthetic personas, friction on industrial-scale inauthentic behaviour, and the defence of deliberation time in high-stakes decisions, since deception at machine speed only works against judgement that has no time to verify. The final defence is relational. Trust built through continuity, in communities and institutions with a shared history, is the one input that synthetic influence cannot counterfeit at scale, because it is not information. It is a relationship.
7.5 The Personal Terms
The six letters are not abstractions. Each one has a personal edition. Honesty is what you check before sharing and what you refuse to launder in euphemism; it is also the habit of demanding receipts from leaders, which is the cheap-talk rule of Section 7.1 applied to politics. Concealment is fought by refusing to treat secrecy as normal, and by supporting disclosure even when the disclosure is unflattering. Restraint is the engineering habit in miniature: write your threshold down before you build, at work as well as in policy. Speed is what you decline when the tool is unexamined and the stakes are real.
None of this substitutes for the institutional agenda. It simply has the advantage of being available on the day it is read.
The agenda above is not a wish list. It is a set of moves on six terms, ranked by return, tested against four words. Whether the moves are actually being made can be checked, which is the business of the next section.
8. Putting Odds on the Record
What follows is not prophecy. It is the discipline of this piece turned on itself: three defined events, three horizons, and three sets of odds, published in order to be graded. Each estimate assumes the scoring of Section 5.1 persists unless stated otherwise, and each is deliberately coarse. Odds stated as ‘roughly 70 per cent’ mean what they say and no more.
8.1 Three Scenarios, With Odds
The capability-surprise scenario. At least one security incident involving unmonitored or guardrail-stripped models, at the scale of the July 2026 breach or larger, occurs by the end of 2028. The odds: roughly 70 per cent. Driven by the terms: honesty low, velocity high, concealment growing by diffusion. The odds would rise if cheap talk continues to substitute for disclosure and the hidden layer keeps expanding unobserved; they would fall sharply if the reporting regime of Section 7.1 were adopted, since verified disclosure converts surprises into known risks.
The enforced-correction scenario. An AI-linked crisis, meaning a major infrastructure failure, financial convulsion or security incident with visible casualties, triggers a coercive pause or intervention on frontier development before 2030. The odds: roughly 40 per cent. Driven by the terms: ambition, velocity and concealment, which is to say the engine and the entire load between them. The odds would rise if acceleration continues at its present rate with the score of Section 5.1 unchanged; they would fall if restraint arrives by rule first, because pre-committed thresholds convert a crisis-imposed correction into a chosen one. That substitution is the entire case for Section 7.2.
The coordination scenario. At least one binding and verified restraint agreement between major powers, covering either human authority over nuclear launch or mandatory reporting for large training runs, is reached by the end of 2028. The odds: roughly 35 per cent. Driven by the terms: restraint and honesty, and by political pressure of the kind already requested in the summer letters of Section 1.2. The odds would rise with further incidents, which convert requests into demands; they would fall if the ‘hoax’ framing of Section 4.3 becomes the settled language of the major governments. Should the agreement land, expect the binding constraint to shift: honesty becomes the limiting term, and the ranking of Section 5.2 reorders accordingly.
8.2 The Scorecard: Grading the Odds
Three estimates cannot establish calibration in any statistical sense; a sample of three grades honesty, not skill. The value of publishing them is the discipline imposed at the moment of writing. Defined events defeat vagueness. Deadlines defeat postponement. Stated odds defeat the temptation to have implied everything in advance. And the clause on what moves each estimate, up or down, keeps every scenario open to evidence rather than to reinterpretation.
The scorecard is therefore not about vindication. At each horizon, the question is not ‘was the writer right’ but ‘were the odds reasonable, and what has happened to the terms’. Readers who disagree are invited to publish their own odds against the same three events, in the same format; disagreement expressed as quantities is exactly what Section 3.3 asked for.
An equation that preaches honesty must accept correction in public. Where reality moves the odds, this model updates, and the update goes on the record beside the original.
9. Conclusion
The estimates of Section 8 carry expiry dates. One horizon ends in 2028, another in 2030. The future carries no such date. It is unwritten, and nobody holds its calendar.
In the weeks before this piece was completed, the public argument produced a perfect illustration of the difficulty. A forecast circulated that artificial intelligence would bring catastrophe by 2030. The chief executive of Nvidia answered with a round number of his own: ‘0% chance’, he said, that the world ends that way by 2030 Jensen Huang dismisses AI doomsday narratives, Fortune, September 2026. Geoffrey Hinton answered him, observing that the seller of the picks and shovels grades the safety of the gold rush Hinton takes on Huang over zero-risk view, NDTV Profit, September 2026. The incentives run the other way as well: catastrophe forecasting pays its own advocates in attention and funding.
Set the two claims side by side. One schedules the apocalypse; the other empties the calendar. Both are certain about a date, and certainty in this domain correlates suspiciously well with interest. Zero and one are the only two probabilities that a single morning can disprove. This piece declines both claims, in the register of its own error bars. We do not know the date or the hour.
That not-knowing cuts in both directions, and holding it honestly is the final discipline. Against fatalism: catastrophe is not scheduled, and the odds of Section 8 are prices attached to present behaviour, not sentences passed on the future. Against complacency: no clock announces the eve of the test. Because the date is unknown, preparation cannot be timed for the eve. It must already be in place. Ambiguity about timing is not an inconvenience to this argument. It is the argument’s last support.
A margin is written for a load of unknown size, arriving on an unknown date. That is what margins are for, and it is the habit with which this piece opened. The technology of the century arrived with nothing of the kind: no equation, only arguments. What can be calculated has now been calculated. A ratio, four laws, six letters, five weights, a reading, three sets of odds. What cannot be calculated, the date, the hour, the form of the test, is the reason the margin matters. The equation is a civilisation’s factor of safety. The unwritten future is the load. If the decimals are wrong, they are at least written down where they can be corrected, which is more than any confident forecast has offered. This is the margin I would write.
9.1 True Power = Capability × (1 − Use)
The maturity described in this piece is the capacity to walk away from one’s own power. Not the inability to act, and not the refusal of capability, but the sovereign choice not to spend it. Every restraint in Section 7 is this choice written small. Every load term in the equation carries the price of refusing it.
The power to act was never the right to act. That distinction, and not any technology, is what a civilisation carries into the unknown.
9.2 Which term will we move?
The unwritten book has no author and no date. It is written in instalments, by millions of people moving six letters without knowing that they are moving them: what they check before sharing, what they demand of leaders, what they refuse to normalise, what they slow down at work. The score of Section 5.1 was a reading of direction, not destiny. Probability is the shadow of present behaviour, and shadows move when the body moves.
The question in this heading is therefore literal. Not what will happen, because nobody knows. Not what should happen to the world, because nobody commands it. Which term, within your own reach, will you move?
The future is unwritten and its date is not ours to know. The margin is written for exactly that reason.











