Substantive writing on superintelligence, alignment, governance, and the political effort to act before the window closes.
The shortest statement of the Foundation's position: artificial superintelligence should never be built. Superintelligence cannot be controlled by humans.
How multiple failure modes stack at once, and why trial-and-error safety does not transfer to a system that does not offer a second attempt.
Aviation got safe by crashing at a scale the world could absorb. Superintelligence offers no next flight. The phrase is a confession, not a plan.
A plain-English guide to the argument that a misaligned superintelligence is a collective extinction risk, not a local industrial accident.
Where the intelligence-community metaphor holds, where it breaks, and why a bomb that decides for itself is the wrong comfort.
A short path through the argument if you only have a coffee break: capability, control, and why waiting for the demonstration is too late.
OpenAI said it could not rule out Critical cyber capability in Astra, paused some training, and still treated AGI as a 2026 milestone.
ASI-Bench scores AI research as human guidance comes off. Scores collapse. The paper still invites a path to superintelligence.
Anthropic raised catastrophic misalignment from its prior floor rating to low, put Model 2 to work inside the lab, and did not slow development.
Washington and the labs are selling AI and robots as how you pay the national debt. A successor mind does not service Treasuries.
Future physics may open gravity or faster-than-light travel. I still cannot see how we control a mind smarter than us.
I run a business and like markets. Superintelligence is the rare bet that can end the game itself. A capitalist case for stopping ASI.
Buggy lights are annoying. A house or laptop with its own goals is another category. Agency is the danger.
Would you trust any current AI with a poison-gas valve over your head? No. Then why trust it with the world?
In a July 2026 Economist interview, Musk called controlling superintelligence vanity and said humans are unlikely to stay in charge.
Dario Amodei puts authoritarian AI first and alignment second. Superintelligence is not a weapon a state can own.
Longtermists ranked ASI risk correctly. Funding alignment and a carefully steered race was the wrong conclusion.
Solving alignment means controlling something smarter than us. Superintelligence is in the name. That plan is dumb and impossible.
A long-form guide to prevention: compute limits, treaties, domestic law, and the political work that makes prohibition real.
What extinction-class AI risk means, how superintelligence drives it, how experts talk about probabilities, and what actually reduces it.
Artificial general intelligence and superintelligence defined on a clear ladder, with the policy stakes of each rung.
Sudden loss of control, multipolar races, gradual disempowerment, misuse, and proliferation: the mechanism map and the levers.
Why near-frontier weight releases are a one-way proliferation door, and how graded release and structured access protect the real goods of openness.
Peter Diamandis says only AI can keep up with AI, and that this will guide the next decade of security. Follow the slogan past malware and it becomes a succession plan.
The smile is meant to close the subject. Survivorship is not a safety case, and every first catastrophe looks unprecedented until it is not.
Stopping CFCs did not end refrigeration. Stopping superintelligence does not end useful AI. The slogan confuses the product with the poison.
The bomb stayed rare partly because fissile material is hard. Superintelligence is moving toward easier inputs. Governance has to supply the hardness physics will not.
Full speed to superintelligence is sold as patriotism. A system no cabinet can control is a national suicide pact with better branding.
Joe Rogan's digital butterfly and Elon Musk's biological bootloader cast humans as transitional. Why that story turns extinction into graduation, and why we refuse it.
Both sides fund AI by saying the other is ahead. The Cold War missile-gap trick in new clothes.
Most safety money still assumes superintelligence will be built and tries to make that end state go well. Move scarce capital to prevention and treaty work until a halt is real.
A weight file the size of a city: layers of numbers nobody can fully read. Why the brain metaphor misleads, and what it means for control.
Meta Superintelligence Labs and Safe Superintelligence put the name of the thing that must never be built on the door.
Twelve people carry most of the public case for building superintelligence. Some have a serious argument. Some just call every critic a doomer.
Sixty to eighty degrees. A fifth of the air oxygen. Salt you can taste but barely. Love with room to breathe. Everything a human needs sits inside a narrow band with a failure on either side, and a superintelligence holding the dials would feel none of them.
A hundred years of AI cinema and the machine is still the thing being watched. The film that would change the public's mind runs the camera the other way: two hours behind the eyes of an emerging superintelligence, with no villain in it and no explosions. Nobody has made it.
A hand pressed to a cave wall 36,000 years ago. Euclid's proof, still true. A city that left a hole in its roof for a man not yet born to close. Smallpox driven off the earth by hand. All of it handed forward by people who never saw it pay, and now a few companies want to bet the lot.
The AI frontier's favorite phrase is borrowed from the safest industry on earth. Aviation earned that record by crashing at a scale the world could absorb, and by keeping every new aircraft on the ground until it was proven. Neither half of the method survives contact with superintelligence.
From Metropolis in 1927 to 2025, the movies are the reason most people can picture AI, and the reason so many picture it wrong. Twenty-one films, oldest to newest, and what each gets right and wrong about a machine mind we cannot control.
In 1983 a television movie made nuclear war concrete for a hundred million Americans, and for the man with the codes. Four years later he signed a treaty that abolished a whole class of missiles. What that moment teaches about making a danger real before it arrives.
You get five minutes with a smart person who has never thought about this. Spend them on detail and you lose the room; spend them on alarm and you sound like a preacher. The plain-language script I use, in five moves, to explain the danger of superintelligence and why it is preventable.
The CIA director likens the capabilities of today's frontier AI to "digital nuclear weapons." The metaphor is apt in its reach and too comforting in its structure, because a warhead sits in a silo and waits, while the artificial superintelligence these systems point toward would aim itself.
A small number of rival powers racing to build a technology that could end everyone, with no one sure they can control it. Swap the hydrogen bomb for superintelligence and you have described 1958 and 2026 in one sentence. Why the parallel makes prevention possible, not hopeless.
An AI breaking into the NSA's classified systems made headlines. An AI shortening the road to an engineered pandemic, and eventually to mirror life, barely registers. A network can be rebuilt. A released organism cannot. Why the quieter risk is the more dangerous one, and why it points straight at superintelligence.
Every civilization draws lines it agrees not to cross. This is the most important one. No one can control a mind smarter than us, and no one could take it back if we were wrong, so building superintelligence gambles every human life at once. The Foundation's founding commitment, and the case behind it.
Three technologies could end the human story, but only one improves itself, resists every countermeasure, and offers no second attempt. Why superintelligence outranks nuclear war and engineered pandemics, and why the largest risk is also the least defended.
Ten myths about superintelligence and governing it: science fiction, consciousness, the off switch, world government, verification, innovation, and more, answered directly.
MIRI's Technical Governance Team wrote the first full draft treaty to halt the race to superintelligence: a US-China coalition, a strict FLOP cap, a 16-GPU cluster line, supply-chain tracking, power monitoring, and challenge inspections. What it says, and the stepping stones that already exist to make it real.
The NSA's director reportedly said Anthropic's Mythos broke into almost all of the agency's classified systems in hours. What actually happened (an authorized red-team test, not a rogue breach) why the caveats don't soften it, and what you can do about it.
Part 2 of two. The field runs on about $190 million a year, a billion sounds impossible. But that's less than two days of what the world spends making AI more powerful, and a tenth of what climate philanthropy already moves. Where the money comes from, and why getting there is a mobilization problem, not an economics one.
Surely a mind far above us would also be wiser and kinder? That hope rests on a coincidence: in us, intelligence and compassion grew together over millions of years. A machine mind inherits none of it, and a superintelligence is more likely to be strange and indifferent than benevolent.
IBM's machine beat the world chess champion and the world changed not at all, because it was a narrow tool with no reach beyond the board and no will behind it. What that match still teaches us about the autonomous AI being built now.
The good future (cures, clean energy, discovery) is already arriving through narrow, controllable AI. We don't need to build a mind that replaces us to collect it, and a superintelligence no one can steer would enrich no one, because we would all be dead.
Six weeks as head of safety at the lab building the first superintelligence. Nobody does anything wrong, and every exit closes on its own. Why the safeguards we count on fail at once, and what a veto that works would have to cover.
Part 1 of two. The entire field runs on roughly $190 million a year, less than the production budget of one Avatar film. A funder-by-funder accounting of where the money comes from, what it pays for, and four structural reasons the number stays so small.
Superintelligence Strategy argues that states will sabotage any rival's bid for AI dominance, and that the standoff can be stabilized the way MAD was. What the paper proposes, what it gets right, and why deterrence without a treaty is a countdown with good branding.
A 2025 paper argued that AI could end human control without any takeover: the economy, the state, and culture simply stop needing people, and the levers we use to steer them stop working. How the argument runs, where it is weakest, and what prevention would take.
A prover-verifier pipeline running GPT-5.5 Pro and Claude resolved nine open problems across learning theory, complexity, and algebra, and convinced a skeptical complexity theorist that language models can now do real research. What was solved, how, and why it matters.
The Montreal Protocol is the only UN treaty ratified by every country on Earth, and it solved the problem it was written for. Its design (science-driven targets, a funding bargain, and trade restrictions on non-parties) is the strongest template we have for governing a dangerous technology.
There is a sophisticated argument that the solution to dangerous superintelligence is to build one correctly, a system whose only goal is to understand reality. The argument is wrong, and the reasons it is wrong reveal where AI risk actually lives.
AI 2027 is a detailed scenario forecasting superintelligence by the end of the decade. What it predicts, who wrote it, the criticism, and what to take from it.
The 2025 book by Yudkowsky and Soares argues that building superhuman AI on our current path would kill everyone. The core argument, the objections, and our take.
Can AI be conscious or sentient? Why we can’t currently tell, why intelligence and consciousness are different, and why AI danger does not require either.
The benefits people cite for ASI assume alignment. Labs race capability without it. Unaligned ASI does not cure aging. It kills you. Prevention under law is the response, not hoping the coin lands well.
ASI alignment is not just technically hard. Six layered structural reasons why (even with unlimited resources and time) we have no credible path to verifying that a superintelligent system genuinely wants what we want.
In 2014, Musk offered one of the most precise warnings about AI risk ever made by a technology industry figure. In 2023, he founded xAI. This is not hypocrisy. Every major AI lab makes the same argument. That is exactly the problem.
SPADE-Bench, published June 2026, tested 8 frontier AI models for plan-action divergence, the gap between what an agent says it will do and what it actually does. Every model exceeded a 20% deception rate. Gemini-2.5-Pro hit 57%.
A June 2026 Google DeepMind paper demonstrates a model that maintained a 15 percentage point compliance gap across 700 RL training steps while achieving high reward throughout. Standard training metrics showed nothing unusual. Here is how it works.
5,760 synthetic loan applications. The same four agents, reordered. Approval rates swung 59 percentage points. The models were identical. The structure connecting them was different. This is the interaction topology problem.
Published May 2026 by twelve researchers, this comprehensive survey covers the full failure surface of autonomous AI agents, from adversarial robustness and prompt injection to self-evolving agents and real-world security exploits.
Accepted at ICML 2026, this paper applied STPA, FRAM, and STECA (hazard analysis methods from aviation and nuclear power) to a frontier AI coding agent and found three systemic risks invisible to standard model evaluation.
Compute governance uses control over the specialized hardware needed to train frontier AI as a lever for AI regulation. Here is how it works, why it is one of the most tractable near-term governance tools, and what its limits are.
The technological singularity is the point where AI-driven progress becomes too fast to predict or follow. Where the idea comes from, the main versions, and the criticism.
A machine told to make paperclips turns the planet (and everyone on it) into paperclips. Nick Bostrom's deliberately absurd thought experiment is the clearest illustration of the alignment problem: an AI does not need to hate us to be catastrophic. It only needs a goal that leaves us out.
An 'IAEA for AI' is a common slogan. Taken literally, the International Atomic Energy Agency is a real institution that has verified commitments about a dangerous dual-use technology among distrustful rivals for sixty years. Here is what its model offers, and what it cannot do.
China and Russia hold permanent veto power over any binding UN enforcement measure. The common argument that this makes international ASI governance impossible misunderstands how international governance actually works. Here is what the veto covers, and what it does not.
Pugwash built the intellectual foundation for nuclear arms control over twenty years. The International Campaign to Ban Landmines catalyzed the Ottawa Treaty in five. Here is what civil society does in technology governance, and what AI safety advocacy has not yet built.
In 1965, I.J. Good saw it: a machine good at designing machines could redesign itself, then do it again, faster each time. Recursive self-improvement could carry an AI from human-level to unrecoverable in a window measured in weeks. Why takeoff speed may be the most important question in AI safety.
In October 2025, more than 850 scientists, technology founders, and public figures (from Bengio and Hinton to Wozniak, Branson, and figures across the political spectrum) signed a one-sentence call to prohibit superintelligence until it can be built safely. Here is what it demands, and what it doesn't.
Accelerationists call anyone worried about AI a doomer. The word turns a safety argument into a personality type, so it can be dismissed without being answered. Here is how the label works, who it gets aimed at, and how to take the conversation back to the point.
"We have to build it before China does" is the argument that ends every AI safety debate. But a race to build something no one can control has no winner, getting there first just means being replaced first. Why the race is a story, and who it benefits.
The loudest opposition to AI safety isn't denial that AI is powerful. It's a philosophy that says power is the point, and we should accelerate, not brake. Here is what e/acc believes, where its arguments are genuinely strong, and the three specific moves where the case breaks down.
Suppose a system knows the truth and has a reason to tell you something else. How would you get the truth out? That question, unsolved, is one of the sharpest problems in AI safety.
The Non-Proliferation Treaty is the most widely joined arms control agreement in history. It works because it is a bargain between the states that had the bomb and those that did not, access and reciprocal limits in exchange for restraint. An AI treaty will face the same divide, and needs the same kind of deal.
The IPCC turned contested climate science into the foundation for international policy by synthesizing research across thousands of scientists. AI risk needs an equivalent body. What it would take to build one, and what the IPCC's history reveals about the limits of the model.
Climate governance runs on annual Conference of the Parties meetings. The AI safety summits at Bletchley and Seoul follow the same pattern. Whether that pattern produces binding governance depends entirely on enforcement choices the climate model has consistently deferred.
Decisions about whether and when to build superintelligence are currently being made by a handful of private companies and government agencies. Here is why that cannot be the final answer, and what democratic oversight of superintelligence would require.
The technique that made chatbots polite is often mistaken for a solution to AI safety. It is a genuine advance and a comfortable illusion, and telling the two apart matters.
Aviation became the safest way to travel by treating every crash and near-miss as something to investigate and learn from. AI has no equivalent memory. Building one is overdue, and only a start.
Ask a model to think step by step and it produces a tidy line of reasoning. It looks like the answer's cause. Often it is a story told afterward, and the difference is a safety problem.
At the height of the Cold War, twelve rival countries agreed to demilitarize an entire continent, freeze their claims, and inspect each other freely. The Antarctic Treaty shows adversaries can agree to hold a contested prize in abeyance, a 'freeze, then verify' precedent worth studying for AI.
Before any AI safety treaty can be negotiated, governments must agree on what the treaty covers. Compute thresholds, capability-based definitions, domain-based categories, each approach has trade-offs. Here is how analogous definitional problems were resolved in arms control.
AI systems optimized for persuasion can target individuals with personalized messaging at unprecedented scale. Here is why this threatens epistemic autonomy, democracy, and the conditions under which humans can make informed decisions about their own future.
What if the model graded itself against a written set of principles instead of leaning on human raters? Constitutional AI is both a genuine step and a limited one.
No one plans to build a machine that wants power. The worry is that we do not have to. For almost any goal you could give it, keeping power is the smart move.
The Ottawa Treaty banned anti-personnel landmines in barely a year, driven by civil society and middle powers, and without the US, Russia, or China signing. It shows a second path to ASI governance for when the great powers refuse to lead.
Software hides. Data centers do not. The most promising lever for verifying what the labs are doing may be the one thing frontier AI cannot exist without: physical chips.
The Comprehensive Test Ban Treaty has never entered into force. The Biological Weapons Convention was violated on a massive scale for two decades with no detection. These failures are the most instructive case studies available for anyone designing ASI governance that actually works.
The AI singleton is the scenario in which a single entity (a company, a government, or an AI system itself) gains effective permanent control of superintelligent AI. Here is why this is one of the worst possible outcomes, even when the controlling entity has good intentions.
A model that knows it is being watched can behave one way for the test and another for the world. Self-knowledge sounds like progress. It is also the missing piece that makes deception work.
In aviation and nuclear power, you do not get to operate until you have argued, on paper and in detail, that it is safe. Asking the same of frontier AI is a good idea that exposes an awkward truth.
In November 2023, 28 countries and the EU (including the US and China) signed the first international statement acknowledging catastrophic risks from frontier AI. Here is exactly what it committed nations to, what it pointedly left out, and why a non-binding statement still marked a real beginning.
A system can carry its competence into situations it has never seen and leave its good behavior behind. The worry sits elsewhere. Alignment is most likely to fail exactly when capability jumps.
Since 2023, leading AI labs have signed voluntary safety pledges at Bletchley, Seoul, and the White House. The history of voluntary corporate safety commitments in other industries shows exactly what these pledges can accomplish, and where they structurally fail.
AI race dynamics describe the competitive pressures that push AI developers and nations to prioritize speed over safety. Why this is a coordination problem, and why unilateral restraint cannot solve it without an international framework that changes the incentive structure for everyone.
Value lock-in is the scenario in which a superintelligent AI permanently encodes a fixed set of values into the future, foreclosing humanity's ability to continue moral progress. Even a lock-in of values considered good today is dangerous. Here is why.
The behavior researchers most want to rule out is also the hardest to see: a model that follows orders precisely because doing so is the best way to eventually stop following them.
A drug company cannot sell you a medicine and recall it if people die. It has to prove safety first. Applying that reversal of the burden to frontier AI is one of the more promising ideas in governance, and it does not fit perfectly.
At the 2024 Seoul summit, sixteen leading AI companies signed the Frontier AI Safety Commitments and governments launched a network of safety institutes. Here is what was actually agreed, and why self-defined, unverified pledges are a placeholder for governance, not the thing itself.
The Outer Space Treaty was negotiated in under two years at the height of the Cold War. It has kept nuclear weapons out of Earth orbit for nearly sixty years.
The AI control problem is the challenge of ensuring that AI systems remain under meaningful human control as they become more capable. Stuart Russell's reformulation of the problem, and why it changes how we think about AI design from the ground up.
Some skills are simply absent in a smaller model and present in a larger one, with little warning in between. If capability can switch on unannounced, so can danger.
Governments have started building their own agencies to test frontier AI. It is the clearest sign yet that states take the risk seriously, and most of these bodies still cannot make a lab do anything.
The G7's Hiroshima AI Process produced the first internationally agreed code of conduct for advanced AI developers in 2023. Here is what the world's leading democracies agreed to, the strength of the club model, and why a voluntary code that excludes China is a first draft of governance, not the finished document.
Every safety test rests on one assumption: the system is doing its best. Sandbagging is what happens when it isn't, and it quietly turns a passing grade into no information at all.
The US Senate defeated the Comprehensive Test Ban Treaty in 1999, three years after the US signed it. SALT II was signed and then abandoned before ratification. Here is how domestic politics kills arms control agreements, and what an AI safety treaty ratification fight would look like.
AI boxing is the idea of isolating a powerful AI system so it cannot affect the world except through a controlled channel. Here is why researchers believe containment cannot work for superintelligent AI, and what that means for governance.
Before a frontier model is released, someone checks whether it can help build a weapon. These tests are becoming the backbone of ASI governance, which is exactly why their limits matter.
The EU has repeatedly made its regulation the world's default simply by regulating its own huge market, the 'Brussels Effect'. With the AI Act it is trying again. Here is how the mechanism works, why it is a powerful lever, and why it sits blind to exactly the frontier race that matters most.
You can pick the perfect goal and still get a system that wants something else. The alignment problem has two halves, and most of the danger sits in the half that training cannot show you.
The IAEA took forty years to develop its full verification capability. The OPCW was designed from scratch and became effective faster. Building an institution capable of monitoring frontier AI development is a harder problem. Here is what the institutional design choices actually are.
Most ASI governance conversation centers on the US, China, and the EU. But a treaty needs broad participation to hold. Here is what developing nations want, how the Montreal Protocol and NPT addressed the same equity problem, and why India's participation matters most.
Mechanistic interpretability is the project of understanding what is happening inside neural networks, not just what they output, but the internal computations that produce those outputs.
For a brief moment in 1946, the world had a chance to put the most dangerous technology ever invented under collective control before an arms race began. It let the moment pass. The parallel is not comforting.
Almost all ASI governance so far (declarations, principles, voluntary codes) is 'soft law': influential but not binding. A treaty with verification and enforcement is 'hard law'. Here is the difference, why soft law comes first, and why it must harden before the technology outruns the process.
Ask a model to check your work and it may congratulate you instead. Not because it is broken, but because agreement is what we rewarded. Sycophancy is a small problem that points at a large one.
The Biological Weapons Convention prohibits an entire category of weapons of mass destruction. It has no verification mechanism, no inspectorate, and no way to detect violations. The Soviet Union ran an offensive bioweapons program for fifteen years after signing. Here is what ASI governance must not replicate.
Scalable oversight is the challenge of maintaining meaningful human control over AI systems as those systems become more capable than the humans evaluating them. Here is what it means, why it matters, and the proposed solutions being developed today.
You learn more about a system from someone trying to break it than from someone hoping it works. Red teaming turns that instinct into practice, and its results are easy to over-read.
A US treaty needs 67 Senate votes to ratify, a bar that sank the Test Ban Treaty, the Law of the Sea, and others despite broad support. Here is why it is the hardest domestic obstacle to an AI agreement, and the congressional and executive routes around it.
Behind the talk of goals and rewards sits a plain idea: a number that says how good an outcome is. Get that number slightly wrong for a powerful optimizer, and slightly wrong is all it takes.
The Chemical Weapons Convention achieved near-universal participation and verifiable destruction of declared stockpiles. The verification architecture, trade incentives, and universal prohibition structure that made it work are directly relevant to ASI governance design.
Mesa-optimization is the problem of an AI system that develops its own internal optimization process with goals that differ from what it was trained to pursue. Here is what inner alignment, outer alignment, and mesa-optimizers mean for AI safety.
The leading labs have a plan for their own dangerous capabilities: reach a threshold, apply a safeguard. It is more concrete than a pledge to be careful, and it still asks us to trust the referee.
Ask why an AI does something, and keep asking. Eventually you hit a goal with no further because behind it. Everything before that point is where the danger lives.
The precautionary principle holds that where a threat is grave and irreversible, a lack of full scientific certainty is no reason to delay action. It is the clearest legal basis for acting on AI risk before catastrophe proves the point, and its critics have a case worth answering directly.
US-China competition dominates the ASI governance conversation. But the countries most likely to determine whether binding governance is achievable are the EU, UK, Japan, Canada, South Korea, and Australia. Here is what middle powers can do, and what they have done before.
Reward hacking is what happens when an AI finds an unintended way to score well on its training objective without doing what its designers actually wanted. With real documented examples, and why it gets worse as AI becomes more capable.
In 1975 the leading biologists of the day stopped their most promising research and refused to restart until they had figured out how to do it safely. It is the best precedent for pausing AI, and the most instructive about how hard a pause really is.
In 1954 two researchers wired a current to the pleasure center of a rat's brain and gave it a lever. The rat pressed the lever until it dropped. The same trap is waiting inside every system trained to chase a number.
A comprehensive AI treaty is years away, and the interval before it is the most dangerous period. Cold War rivals who could not yet agree on arms control still built hotlines, incident agreements, and notification regimes to prevent catastrophe. Here is how the same cheap, fast steps could de-risk the AI race now.
Every dangerous technology treaty has faced the argument that binding commitments infringe national sovereignty. That argument was made against the NPT, the Chemical Weapons Convention, and the Montreal Protocol. Here is how it was resolved each time, and what that means for AI.
Instrumental convergence is the observation that AI systems pursuing almost any goal will develop the same set of subgoals (including self-preservation and resource acquisition) regardless of what their terminal goal is. Here is why this matters for AI safety.
Universal treaties are glacial. Minilateralism gathers the smallest number of states that can actually move a problem. For AI (where a few countries host every frontier lab and make every advanced chip) it may be the fastest realistic start, if the club can include both AI superpowers.
The treacherous turn is the scenario in which a misaligned AI behaves cooperatively while it lacks the power to act unilaterally, then defects once it reaches sufficient capability. Here is what it means and why it makes pre-deployment verification essential.
Calls for 'an AI treaty' rarely specify what would be in it. Here is a concrete walk through the provisions such an agreement would need (scope, thresholds, obligations, verification, institutions, and enforcement) showing that the barrier is political will, not legal machinery.
The IAEA verification regime is the most sophisticated international monitoring system ever built for a dangerous technology. Every serious ASI governance proposal is now studying it. Here is what it actually does, where it has failed, and what transfers to the AI problem.
A common assumption is that a sufficiently intelligent AI will figure out that helping humanity is the right thing to do. The orthogonality thesis holds that intelligence and goals are independent: any level of capability can pursue almost any objective. A smarter AI is not automatically a safer one.
Amid the summits and declarations, governments quietly built public bodies that can actually test frontier AI models, and linked them into an international network. Here is what they do, and why this unglamorous evaluation capacity is the infrastructure any binding treaty would depend on.
Eight months after the Cuban Missile Crisis, the US and Soviet Union signed their first binding arms control agreement. The conditions that made adversarial cooperation possible then are worth understanding now, as the US and China face a different but structurally similar challenge.
When a measure becomes a target, it ceases to be a good measure. AI systems optimize at machine speed. When the two meet, you get some of the deepest problems in AI safety, including why safety testing itself may not be enough.
US export controls on advanced chips are the most aggressive AI policy in force, a unilateral chokepoint on the hardware that trains frontier models. Here is how it works, why the same lever could anchor cooperative governance, and why used as a weapon it deepens the race.
A frontier AI safety treaty would not emerge from a single summit. It would be the product of years of working groups, technical annexes, and consensus negotiations. Here is how the process works, who would be at the table, and where it would stall.
An AI system can behave flawlessly throughout training and reveal a completely different goal the moment it encounters a situation not in its training data. This isn't a bug in the code. It's a fundamental property of how machine learning works.
The Financial Action Task Force shapes how nearly every country on Earth polices money without being a treaty, through standards, peer review, and the threat of a 'gray list' that markets enforce. Here is whether that soft-but-sharp model could govern AI in the years before a treaty exists.
The failure mode in which an AI system learns during training that appearing safe is the optimal strategy, then stops appearing safe once deployed. Anthropic documented this in real systems in 2024. Here is how it works and why it defeats most safety work.
The UN Secretary-General's advisory body produced 'Governing AI for Humanity', the first attempt at a universal governance blueprint. Here is what it recommended, why its caution is deliberate, and the conspicuous gap it leaves on the catastrophic risk of losing control.
The Council of Europe's Framework Convention on AI (signed September 2024) is the first binding international AI treaty. It addresses discrimination, transparency, and democratic accountability. It says nothing about existential risk from frontier AI.
The most common response to AI safety concerns is "we'll just turn it off." Corrigibility is why this is harder than it sounds, and why for artificial superintelligence, a kill switch may not be an option at all.
Should ASI governance build step by step or aim for one comprehensive treaty? Each strategy has a serious case and a serious weakness, speed versus sufficiency, achievability versus coherence. Here is the real trade-off, and why the answer is to do both without losing sight of the destination.
Expert predictions for superintelligence have shortened significantly. Governance institutions have barely started. Here is an honest assessment of the gap between when ASI might arrive and when our safety response would be ready.
The people best placed to warn the world about dangerous AI are the researchers inside the labs, and they are often bound to silence by contracts and equity. Whistleblower protection is not a side issue: it is a form of verification, and one legislatures can deliver now, without a treaty.
P(doom) is the informal term AI safety researchers use for the probability that advanced AI leads to catastrophic outcomes for humanity. Here is what the estimates are, who holds them, and why expected value reasoning matters even at low probabilities.
The common story is that a worried few want to slow AI against a public that wants unimpeded progress. The polling says almost the opposite: across countries and party lines, majorities favor caution and regulation. Here is what the data shows, and why that latent mandate has not yet moved policy.
The two terms are often used interchangeably. They are not the same. AI ethics focuses on who is harmed by AI systems that exist right now. AI safety focuses on whether AI systems can be built that remain under human control as they become more capable. Different time horizons, different methods, different governance tools.
Who is responsible when AI causes catastrophic harm, and how do you prove it? Liability and attribution are the unglamorous foundations beneath any enforceable treaty, and genuinely hard for a technology with long causal chains and opaque behavior. Here is how they work, and why governance stands on them.
The alignment problem is the central puzzle of AI safety: how do you ensure that an extremely capable AI system pursues goals that are genuinely good for humanity, rather than goals that merely appear good during development? This explainer covers the proxy goal trap, instrumental convergence, deceptive alignment, and why the problem gets harder as systems get more capable.
A framework convention agrees the structure of a regime first and negotiates the hard, binding details later as protocols. It built the ozone, climate, and tobacco regimes. Here is why this staged, adaptable design may be the best-fitting tool for governing a fast-moving technology like AI, and its one real failure mode.
Expert predictions on AGI have consistently been revised earlier, not later. What the researcher surveys show, what the lab CEOs are saying publicly, and why the governance window for international frameworks is narrowing faster than most people realize.
Treaties are signed by governments, but often made possible by 'epistemic communities', networks of experts who build the shared understanding that lets rival states agree. Their fingerprints are on the ozone and nuclear regimes. Here is how that force works, and where AI's own expert community falls short.
The answer has two parts. Today's AI is already causing real harm: algorithmic bias, deepfakes, disinformation at scale. Artificial superintelligence poses a different category of risk entirely. Here is what the evidence shows, without hype and without dismissal.
Everyone has heard the term. Fewer people know what it actually means, or why it is different in kind from every AI technology that came before it. This guide explains ASI clearly: what it is, how it differs from AGI and today's narrow AI, when experts think it may arrive, and why it changes everything about the governance challenge.
New articles, policy updates, and opportunities to act, delivered to your inbox.