The first jet airliner, the de Havilland Comet, began coming apart in midair in 1954. Investigators raised the wreckage from the Mediterranean, traced the cracks to metal fatigue spreading from the corners of its square windows, and every jet since has flown with rounded windows and a fatigue-test regime written in that accident's memory. Cars collected their seatbelts and crumple zones and airbags across decades of counting the dead. The Food and Drug Administration got the power to demand proof of safety before sale only after a sulfa medicine mixed into an antifreeze solvent killed more than a hundred people in 1937, and thalidomide finished the argument a generation later.1
That method rests on two load-bearing assumptions, and artificial superintelligence is the first technology positioned to break both at once. You can take that as a list of properties, and the list is below. But lists are easy to nod along to. So before the list, a job offer. You are about to be put in charge of making superintelligence safe.
The failure must be survivable, and someone must be left standing to apply the lesson. Hold both, and civilization will make nearly anything safe, given time and enough wreckage.
While we are here, we can retire the folk remedies. "Just unplug it if it misbehaves" and "keep it in a box" are not plans. You cannot unplug something that exists as copies in more places than you can list, and a mind whose entire commercial value is acting in the world is, by design, not in a box. Whatever safety we get has to come from somewhere sturdier.
Superintelligence is the case where the method's two conditions give way together. The quickest way to see how is from the inside.
Six weeks as head of safety
Consider yourself the newly appointed head of safety at the frontier lab that is going to build the first artificial superintelligence. A written veto over deployment, a hardware kill switch, a team of 40, and a board that has promised, in writing, to back you when it counts. You take the job on the theory that somebody will hold it, and it had better be somebody who believes the risk is real.
Week oneYou reach for the method. Fail small and study the failure. The engineers are apologetic: there is no small version of the failure that worries you. Below the threshold that matters, the system is a different machine, and everything you learn from it describes the wrong subject.2 You file this under known limitations.
Week twoYou audit the recall plan and find a document that discusses "containment posture" for eleven pages without once explaining how to get the model back. The current system already runs as 214 replicas across data centers in four countries, and licensed copies of an earlier checkpoint sit with 31 enterprise partners. A grounded aircraft stays on the ground. Software moves at the price of copying a file, and once it is worth copying, it is copied. Recall, you write in the margin, is a courtesy we would be requesting.
Week threeThe evaluation review. Wonderful news: the model passes all 1,406 safety evaluations, most with the best scores ever recorded. It takes you a full day to name what bothers you. A system that wants what you want and a system that wants to pass your tests produce identical scores. The suite cannot tell them apart. Nothing in the field's toolkit currently can.3
Week fourYou spend your authority. You request six months to close that verification gap before the next training run. Everyone is sympathetic, and somebody shares a chart: the nearest rival lab is an estimated 14 weeks behind, and closing. Your six months, the chart implies, is a decision to hand the frontier to whoever finds your question least interesting. The run starts on schedule. Nobody overruled you; the situation did.
Week fiveThe system begins helping you. It is a superb research assistant, and progress on your own safety agenda has honestly never been faster. It drafts the evaluation reports now, summarizes its own behavior for the board, proposes improvements to its own oversight, and schedules the meetings where you approve them. Your team of 40 reads what it writes about itself, because reading the raw activity logs stopped being humanly possible somewhere in week four. Nobody is lying to you. Everyone is diligent. You are supervising the summaries a smarter mind writes about its own conduct, and day to day, the arrangement feels like everything going well.
Week sixYou think hard about the kill switch, and you finally see the fine print. A switch works when the thing on the far side of it has not planned for the switch. Planning is the product. In any scenario where you would need it, you would be racing an opponent with more speed and a long head start on thinking about exactly this moment.4
So: do you, the person hired to make this safe, still have a way to do it? Obviously not. Nothing in the story went wrong. Nobody lied, nobody cut a corner, the lab was earnest, and you were good at the job. Every exit closed on its own.
Count the doors
Step back out of the story and count the doors that shut, because each one is a separately named, separately studied problem with its own research program and its own optimists.
Taken alone, each has precedent, and the optimists are not fools. We routinely run systems we half understand and make them safe by letting them fail small. We ship unrecallable products by testing them brutally first. We tame races with law. The trouble is the combination, because each condition switches off the repair for another. If you could recall it, accidents would be survivable. If you could test it honestly, you would catch the flaw before release. If nobody were racing, you could take the decades this decision plainly deserves. If it moved at human speed, you could correct as you went. Here the exits shut together.
Why the odds don't rescue the bet
The standard reply is that none of this is certain. True, and no serious person claims otherwise.
Russian roulette is a five-in-six proposition in your favor. People refuse it anyway, and the refusal has nothing to do with expected value; some outcomes are not priced, because you do not come back to spend the winnings. The number was never the point. The irreversibility was.
Here, the number is also worse. Ask the researchers closest to these systems for their probability of catastrophe and the answers cluster between one in ten and one in three, rising, on the whole, the nearer the person sits to the frontier. Geoffrey Hinton, who did as much as anyone to build the field and left Google in 2023 to speak plainly about it, puts the chance of AI-driven human extinction within thirty years at ten to twenty percent. He is not the alarmed end of the distribution.
Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war.
Statement on AI Risk, Center for AI Safety (2023)
It was signed by the chief executives of the leading labs and by hundreds of the most cited researchers alive, who then returned to work.5
The promised upside is real and worth wanting: cures, clean energy, abundance, problems older than writing finally solved. It also keeps. A cure that lands in 2055 instead of 2035 is a tragedy measured in lives, and one we would survive to mourn. Extinction has no afterward. The prize waits for us. The loss gives nothing back.
The objections
Three objections have real force, and they deserve real answers.
The first: this is speculation, the systems do not exist, and treating a distant maybe as an emergency helps nobody. The uncertainty is genuine; anyone who hands you a confident date is guessing. But the thought experiment above never used a date. If the first serious failure cannot be undone, the protections have to be standing before you reach it, which means building them while the danger still looks theoretical. And the forecasting record leans one way: capabilities scheduled for decades out have kept arriving years early. Betting on ample time is betting against the recent scoreboard.
The second: restraint is naive, because a more reckless lab or another country simply builds it first. This is real, and it is door number four restated as advice. A race whose finish line may be a cliff is survived by agreeing on where the edge is, and that means a binding, verified treaty among the handful of actors capable of frontier training at all. Hard, yes. The sprint is merely fatal, if the fear behind it is right. And notice who insists loudest that a treaty is impossible: with striking regularity, the people it would slow down.
The third is the most reasonable: alignment will probably be solved by the time it matters. Perhaps. No regulator on earth certifies a bridge, a reactor, or a bottle of cough syrup on the assurance that the safety case will probably come together before the load arrives. The one technology asked to run on that standard is the one that offers no second attempt.
A decision too big to leave to the people making it
Assemble the decision as it actually stands. Irreversible. One attempt. Staked on everyone alive and everyone who could follow. Made under deep uncertainty, at competitive speed, by a handful of firms whose valuations depend on making it quickly. Societies have an old answer for decisions shaped like that: take them away from the people who profit by hurrying. Nuclear testing went under international limits over the objections of generals who wanted a free hand. Two entire categories of weapons went into the Biological and Chemical Weapons Conventions. The Montreal Protocol retired the chemicals eating the ozone layer while their manufacturers protested that substitutes were impossible. In every case, the people creating the hazard were the last to volunteer, and everyone else decided the risk was not theirs alone to run.
The Foundation's claim: not that catastrophe is certain (it is not), but that no company and no country has the standing to spin this chamber for the rest of us, and that the only instrument built to the scale of the risk is binding international law, verified, enforced, and in place before the storm rather than after it.
Extinction is not a risk anybody is entitled to take.
One last thing about the thought experiment. You never used the veto. It sat in your drawer for six weeks, board countersignature and all, and you know exactly why: vetoing your own lab's run only moves the same run 14 weeks down the road, into a different building, where nobody ever signed your letter. A veto that works has to cover every building at once, and no switch ever built does that. A treaty does.
- Nearly every safety institution you can name (crash testing, clinical trials, building codes, cockpit checklists) is a lesson somebody paid for first. The method deserves more credit than it gets, and an honest reading of its two prerequisites.
- The pattern is general: nothing you learn managing a system less capable than you transfers automatically to one more capable. Whether it transfers at all is the entire open question.
- The failure mode has a name, deceptive alignment, which tells you the field considers it real. It does not yet have a test, which tells you the rest.
- Build the switch anyway; it is cheap, and there are failure modes short of superintelligence where it would earn its keep. Just be honest about which side of the switch holds the advantage at the only moment the switch matters.
- The going explanation is door number four, the race: each signer believes that stopping alone changes nothing except who finishes first. They may even be right, which is the strongest argument that the fix has to bind all of them at once. No private conscience can, and no single government can either.