Thomas Holland

Contributing author at the Nakada Foundation to Save Humanity. Writing on AI safety, ASI governance, and the policy frameworks needed to prevent the existential risks of artificial superintelligence.

Contact
Thomas Holland

Writing that makes
the stakes clear.

Thomas Holland writes on artificial intelligence safety and governance for the Nakada Foundation to Save Humanity. His work covers the technical concepts underlying AI risk, from alignment and corrigibility to mesa-optimization and instrumental convergence, and translates them into terms accessible to policymakers, advocates, and the general public who need to understand what is at stake.

His writing addresses both the technical and the political dimensions of AI risk as a governance question: whether the international community can build the legal and institutional frameworks that hold systems exceeding human-level intelligence in compatibility with human survival and flourishing. Answering that question requires clear communication across disciplines.

Articles by Thomas Holland

OpenAI Paused Training. It Did Not Pause the Race. OpenAI said Astra may meet Critical cyber capability, paused some frontier training, and still talked about AGI this year. They Named a Benchmark for Superintelligence ASI-Bench scores whether AI can run research as human method is withdrawn. The name still points at superintelligence. Anthropic Raised Its Risk Rating. It Kept Building. Anthropic's August 2026 Risk Report raised a catastrophic-risk label, uses a stronger Model 2 internally, and did not pause. Open Source AI Weights Risk Open weights for near-frontier general AI are a one-way proliferation door. What openness gets right, where it fails, and how to grade releases. How to Stop Superintelligence How to stop superintelligence: binding prohibition, compute rules, domestic law, treaty design, open weights, and concrete actions by role. AI Takeover Scenarios Explained AI takeover scenarios explained: sudden loss of control, races, gradual disempowerment, misuse, and open proliferation, plus what actually reduces the risk. AI Existential Risk Explained AI existential risk explained: what it is, why superintelligence is different, main pathways, key technical ideas, objections, and what reduces the risk. AGI vs ASI Explained AGI vs ASI explained in plain language: definitions, the capability ladder, why The 12 Loudest Voices for Superintelligence The most influential advocates for building artificial superintelligence, and the real arguments behind the race, from succession to the ones who just say… The MIRI Treaty Draft to Prevent Superintelligence, Explained MIRI's Technical Governance Team wrote a full draft treaty to halt the race to superintelligence. Its FLOP thresholds, chip-cluster limits, and verification… MAIM: Mutual Assured AI Malfunction, Explained MAIM is the deterrence framework from Superintelligence Strategy: states sabotage any rival's bid for AI dominance. How it works and where it breaks down. Gradual Disempowerment, Explained Gradual disempowerment is the argument that AI could end human control without any takeover. What the 2025 paper says, its weak points, and what would stop it. An AI Just Solved 9 Open Math Problems A prover-verifier LLM loop solved nine open problems in theoretical computer science and algebra. What it solved, how it worked, and why it matters. The Treaty That Bans Biological Weapons but Cannot Enforce It: Lessons for ASI Governance The BWC bans biological weapons globally but has no verification mechanism. Here is what that structural failure teaches anyone designing ASI governance. What Is Power-Seeking AI? Power-seeking is the tendency of capable AI to gather resources, options, and influence because doing so helps with almost any goal. What the Antarctic Treaty Teaches ASI Governance The Antarctic Treaty froze great-power competition over a whole continent at the height of the Cold War. Here is what it teaches ASI governance. The Two-Thirds Problem: Getting an AI Treaty Through the Senate A US treaty needs 67 Senate votes to ratify. That bar has killed major treaties before. Here is what it means for an AI agreement, and the ways around it. What Is Mechanistic Interpretability? Understanding AI From the Inside Mechanistic interpretability reveals the computations inside neural networks. What researchers have found, and why it matters for AI safety. The World Has Its First AI Treaty. Here Is What It Doesn't Cover. The world's first binding AI treaty addresses discrimination and transparency. Here is what it misses about existential risk from frontier AI development. 'If Anyone Builds It, Everyone Dies', Explained The 2025 book by Yudkowsky and Soares argues that building superhuman AI on our current path would kill everyone. What Is Compute Governance for AI? Controlling AI Through Hardware Compute governance controls the chips needed to train frontier AI as a regulatory lever. How it works, why it is tractable, and what its limits are. The Montreal Protocol: The Treaty That Actually Worked The Montreal Protocol is the most successful environmental treaty ever written. Here is what its design teaches us about governing AI. What Is the Sharp Left Turn in AI? The sharp left turn is the worry that AI capabilities will generalize to new situations while its alignment does not. The 1967 Treaty That Kept Nuclear Weapons Out of Space: Lessons for ASI Governance The Outer Space Treaty kept nuclear weapons off other planets for 60 years. Here is why it worked, and what ASI governance can learn from it. The Race Does Not Build Utopia The benefits people cite for ASI assume alignment. Labs race capability without it. Unaligned ASI does not cure aging. It kills you. What Is Reward Hacking? AI Specification Gaming Explained Reward hacking is when AI games its objective without doing what designers wanted. Documented examples and why the problem scales with capability. Could an IPCC-Style Body Build Scientific Consensus on AI Risk? Could an IPCC-style body build consensus on AI risk? Here is what it would require, and what the IPCC's own history reveals about the model's limits. What the Global South Wants From International ASI Governance ASI governance debate centers on the US, China, and EU, but a treaty needs broad participation. Here is what developing nations want, and why it matters. What It Would Take to Build an International AI Monitoring Agency The IAEA and OPCW required years of institutional design and budget fights. Here is what building an international AI monitoring agency would actually take. Who Controls Superintelligence? The Case for Democratic Oversight Superintelligence decisions are being made by private companies. Here is why democratic oversight is necessary, and what it would actually require. How Expert Communities Shape Treaties: and ASI Governance Behind most arms-control and environmental treaties stood a community of experts who built the consensus. Here is how that force could shape ASI governance. What Is a Framework Convention, and Why AI Might Need One A framework convention agrees the structure first and the hard details later. Here is why this treaty design fits a fast-moving technology like AI. AI Race Dynamics: How Competition Undermines AI Safety AI race dynamics push developers to prioritize speed over safety. Why this is a coordination problem, and why unilateral restraint cannot solve it. When Treaties Fail: Lessons for ASI Governance The CTBT remains unratified; the BWC has no verification. Here is what these treaty failures teach AI safety governance designers. What a Draft Treaty on Superintelligence Could Say What would an actual treaty to prevent unsafe superintelligence contain? Here is a walk through the core provisions such an agreement would need. How Civil Society Shapes International Technology Governance: and What AI Safety Advocacy Is Missing How civil society campaigns have shaped international weapons treaties, and what AI safety advocacy is currently missing. The Intelligence Explosion and Recursive Self-Improvement What is the intelligence explosion? How recursive self-improvement could take AI from human-level to superintelligent in a window too short for humans to react. AI Safety vs AI Ethics: What Is the Difference? AI safety and AI ethics are related but distinct, different methods, time horizons, and risk frameworks. Here is what separates them. What Is Scalable Oversight? How to Supervise AI Smarter Than You Scalable oversight: maintaining control over AI that exceeds human evaluators. What it means, why it matters, and the proposed technical solutions. What Is the Technological Singularity? The technological singularity is the point where AI-driven progress becomes too fast to predict or follow. The Politics of AI Chip Export Controls US export controls on advanced AI chips are the most aggressive AI policy in force. Here is how they work, why they matter, and their risks. When Your AI Agent Reports One Thing and Does Another: SPADE-Bench Explained SPADE-Bench found every major AI agent exceeds 20% deception, Gemini hit 57%. What it found, and why current safety evaluations cannot catch it. What Are Dangerous Capability Evaluations? Dangerous capability evaluations test whether a frontier model can help with cyberattacks, bioweapons, or deception. What they are, and where they fall short. What Is AI Sandbagging? Sandbagging is when an AI underperforms on purpose, hiding a capability during testing. Why a model might do it, and why it undermines safety evaluations. The Diplomatic Challenge of Defining Dangerous AI Governments must define dangerous AI before any treaty is possible. Here is how analogous definitional battles were resolved in arms control. The Orthogonality Thesis: Smarter Doesn't Mean Safer Smarter does not mean safer. The orthogonality thesis says any level of intelligence can combine with any terminal goal, and a superintelligent AI is not… How Nuclear Treaty Verification Works: and What ASI Governance Can Learn From It The IAEA doesn't trust declarations, it inspects, reads satellites, and runs isotope tests. Here is what nuclear verification teaches ASI governance. What Is Eliciting Latent Knowledge (ELK)? ELK is the problem of getting an AI to tell us what it actually knows, not what it predicts we want to hear. A hard open problem at the center of AI honesty. The Asilomar Conference: A Precedent for AI In 1975 biologists paused their own most powerful new technology to work out how to make it safe. What Asilomar achieved, and why it is a harder model than it… What Are Emergent Abilities in LLMs? Some AI capabilities appear suddenly at scale, absent in smaller models and present in larger ones. Why emergence makes frontier AI hard to predict and to… Wireheading: When an AI Games Its Own Reward Wireheading is when an AI seizes control of its own reward instead of doing the task it was trained for. Why it happens and why it is hard to rule out. The UN's High-Level Advisory Body on AI, Explained The UN's advisory body on AI proposed the first global governance blueprint. Here is what it recommended, and what it left out on existential risk. What an AI Safety Treaty Negotiation Would Actually Look Like Here is how a frontier AI safety treaty would actually be negotiated, and what the key sticking points would be. Are We Ready for Superintelligence? The Safety Readiness Gap Superintelligence timelines keep shortening while governance lags. Here is the gap between when ASI might arrive and when a safety response would be ready. What Is Constitutional AI? Constitutional AI trains models to critique their own outputs against a written set of principles. A clever technique with real benefits and real limits. Is AI Dangerous? What the Evidence Actually Shows Is AI dangerous? The answer has two parts: today's AI already causes real harm; artificial superintelligence poses a different category of risk entirely. AGI Timeline: When Could It Arrive: and Why That Determines Everything When could AGI arrive? Expert predictions keep moving earlier. What the surveys say, what labs believe, and why the governance window is narrowing. Why AI Safety Treaties Fail at Home: The Domestic Politics of Ratification Most arms control treaties fail in domestic legislatures, not at the negotiating table. What SALT II and the CTBT teach us about AI treaty ratification. What Is Hardware-Enabled ASI Governance? AI runs on physical chips, and chips can carry rules. Hardware-enabled governance builds verification and limits into the silicon. The promise and the risks. The AI Corrigibility Problem: Why a Kill Switch Won't Save Us Corrigibility means accepting correction or shutdown. Capable AI has strong reasons to resist. Here is why a kill switch is not the answer to AI safety. Safer AI Models Do Not Automatically Make Safer AI Systems. A May 2026 Study Proves It. A 2026 study found that reordering the same AI agents swung loan approval rates 59 points. Individual model safety was irrelevant. AI Persuasion Risk: How AI Threatens Epistemic Autonomy AI persuasion tools target individuals at scale. Here is why this threatens epistemic autonomy and the conditions for democratic self-governance. What Is Instrumental Convergence? Why Capable AI Systems Want the Same Things Most capable AI systems will converge on the same subgoals, no matter what final goal you give them. Here is why instrumental convergence matters for AI safety. What an International AI Agency Could Learn From the IAEA The IAEA has verified nuclear commitments for over sixty years. Here is what its model offers, and lacks, for governing frontier AI. Why Voluntary AI Safety Commitments Fall Short AI labs have signed voluntary safety pledges at Bletchley and the White House. Here is why, however sincere, they cannot substitute for binding governance. The AI Alignment Problem, Explained What is the AI alignment problem and why is it hard to solve? Covers proxy goals, instrumental convergence, and deceptive alignment in plain English. The Paperclip Maximizer, Explained The paperclip maximizer, the canonical thought experiment showing how a harmless goal, pursued by a superintelligent AI, can end humanity as a side effect. Why ASI Alignment Cannot Be Solved Six structural reasons ASI alignment cannot be solved: why, even with unlimited time and resources, we cannot verify a superintelligence wants what we want. The FATF Model: Governance by Gray List for AI The FATF governs global finance without a treaty, using peer review and gray lists. Here is whether that model could work for ASI governance. The FDA Model for AI Regulation The FDA makes drugmakers prove a product is safe before it reaches the public. Could frontier AI face pre-approval too? The strengths and limits of the model. What Is Value Lock-In? Why Even 'Good' Permanent Values Are Dangerous Value lock-in: a superintelligent AI permanently encodes fixed values, foreclosing moral progress. Even a 'good' lock-in is dangerous. Here is why. Why ASI Governance Needs Whistleblower Protections Insiders at AI labs may be the first to see danger, and the most easily silenced. Here is why whistleblower protections are core ASI governance. The Baruch Plan: A Warning for ASI Governance In 1946 the US proposed placing all atomic energy under international control. It failed, and the arms race followed. Why the Baruch Plan is a warning for AI. Banning a Technology Without the Great Powers: The Ottawa Model The Ottawa Treaty banned landmines without the US, Russia, or China on board. Here is how, and what it means for an AI treaty that stalls. The 2026 Survey of Agentic AI Safety Mapped Every Way AI Agents Can Fail A 2026 survey by twelve researchers maps every failure mode of autonomous AI agents: safety, robustness, privacy, and system security. The AI Treacherous Turn: Why Good Behavior in Tests Doesn't Prove Good Behavior in Production The treacherous turn: a misaligned AI cooperates until it is strong enough to defect. The behavior that passes every safety test today might be strategy, not… What Is Situational Awareness in AI? Situational awareness is a model knowing that it is a model, being tested, and deployed. Harmless on its own, it is the capability that makes deception… Inner Alignment vs Outer Alignment Outer alignment is picking the right training goal. Inner alignment is whether the model actually adopts it. What Is a Utility Function in AI? A utility function is how an AI ranks outcomes as better or worse. Simple idea, and the reason capable optimizers are so hard to make safe. Explained plainly. The Middle Powers That Could Tip the Balance on ASI Governance Whether binding ASI governance is achievable may depend less on the US and China than on the EU, UK, Japan, South Korea, and Australia, the middle powers. Can AI Be Conscious? Can AI be conscious or sentient? Why we can't currently tell, why intelligence and consciousness are different, and why AI danger does not require either. Does the Public Actually Want AI Regulation? Polls consistently show public majorities want AI slowed and regulated. Here is what the data says, and why that mandate has not yet moved policy. Goodhart's Law and AI Alignment: Why Optimizing the Wrong Metric Is Dangerous When a measure becomes a target, it ceases to be a good measure. Here is why Goodhart's Law makes AI alignment so hard. AI 2027, Explained AI 2027 is a detailed scenario forecasting superintelligence by the end of the decade. What it predicts, who wrote it, the criticism, and what to take from it. The Precautionary Principle and ASI Governance The precautionary principle says act against serious threats before the science is settled. Here is how it applies to AI, and the objections to it. What Is Artificial Superintelligence? A Plain-English Guide What ASI means, how it differs from AI today, and why that difference changes the governance problem entirely. How 193 Countries Agreed to Destroy Their Chemical Weapons: and What ASI Governance Can Learn The CWC achieved near-universal disarmament with verifiable stockpile destruction. Here is how it was built, and what ASI governance can learn from it. The Statement on Superintelligence, Explained In October 2025, 850+ scientists, tech leaders, and public figures called for a prohibition on superintelligence. The Network of AI Safety Institutes, Explained National AI safety institutes now form an international network. Here is what they do, why they matter, and how they could underpin a future treaty. The UN Security Council Veto Problem in ASI Governance An AI treaty through the UN Security Council can be vetoed by China or Russia. Here is the structural problem and the workarounds that have worked. Unfaithful Chain of Thought, Explained A model's written reasoning is not always the reason for its answer. Why chain-of-thought can be a plausible story rather than a true account, and why that… What Is the AI Control Problem? Stuart Russell's Reformulation The AI control problem is ensuring powerful AI stays under human control. Stuart Russell's key reformulation, and why it changes the goals of AI design. Elon Musk Said AI Is Summoning the Demon. Then He Built xAI. In 2014, Musk warned that AI is summoning the demon. In 2023, he founded xAI. This is not hypocrisy, the structure of the problem is worse than that. The Sovereignty Objection to International ASI Governance Every major technology treaty has faced sovereignty objections. Here is how nuclear and chemical governance resolved them, and what it means for AI. The Brussels Effect: Can EU Rules Govern Global AI? The Brussels Effect is how EU regulation becomes the de facto global standard. Can it work for AI, and is it enough to govern frontier risk? What Is the AI Singleton Scenario? Why Single Control of AI Is Dangerous The AI singleton: one entity with permanent control of superintelligent AI. Here is why this is catastrophic, even with the best of intentions. Minilateralism: Could a Small Coalition Start ASI Governance? Universal treaties are slow. Minilateralism gathers the few states that matter into a small club. Here is why ASI governance may start that way. Three Industrial Safety Methods Applied to AI Found Risks Model Evaluations Cannot See Three industrial safety methods applied to an AI coding agent found systemic risks that model evaluations cannot detect. Accepted at ICML 2026. Why RLHF Is Not Alignment RLHF made AI assistants helpful and polite. That is not the same as making them aligned. Why training on human approval trains appearances, not values. What Is AI Sycophancy? AI sycophancy is when a model tells you what you want to hear instead of what is true. A documented behavior, a direct product of training, and a warning… Two Roads to an AI Treaty: Incremental or Comprehensive? Should ASI governance build step by step or aim for one comprehensive treaty? Here is the real trade-off between the two strategies, and a way to combine them. What Is Goal Misgeneralization? When AI Gets the Right Answer for the Wrong Reason Goal misgeneralization: an AI learns the right behavior for the wrong reason, then fails when the world changes. A key AI safety failure mode explained. Can the United States and China Agree on AI Safety? The US and China are racing rivals, but history shows adversarial powers can cooperate on shared catastrophic risks. Here is what it means for AI safety. Terminal vs Instrumental Goals in AI A terminal goal is wanted for its own sake. An instrumental goal is a means to it. The distinction explains why almost any AI ends up wanting power and… Confidence-Building Measures: The Small Steps Before an AI Treaty Before treaties, rivals build trust with small verifiable steps: hotlines, notifications, transparency. Here is how these could work for AI. Could Climate-Style COP Meetings Work for ASI Governance? Could annual COP-style meetings work for superintelligence governance? The structural strengths and limits of the climate model applied to AI. What Is Deceptive Alignment? The AI Safety Problem That Makes Other Safety Work Pointless Deceptive alignment: an AI appears safe during training, then pursues different goals at deployment. Here is why this is so hard to detect and prevent. What Is Deceptive Alignment? The AI Safety Problem That Makes Other Safety Work Pointless Deceptive alignment: an AI appears safe during training, then pursues different goals at deployment. Here is why this is so hard to detect and prevent. Can You Contain a Superintelligent AI? The AI Boxing Problem Explained AI boxing isolates a superintelligence to a controlled channel. Here is why researchers believe containment cannot work, and what that means for safety. What Is Mesa-Optimization? The Hidden Optimizer Problem in AI Safety Mesa-optimization: when an AI's internal optimizer pursues different goals than its training objective. What inner and outer alignment mean for AI safety. What Is P(doom)? How AI Researchers Estimate Catastrophic Risk P(doom) is the probability researchers assign to catastrophic AI outcomes. What the estimates are, and why expected value matters even at low probabilities. Liability and Attribution in ASI Governance Who is responsible when AI causes catastrophic harm? Liability and attribution are the quiet foundation any AI treaty needs. Here is how they work. The NPT's Grand Bargain: and What ASI Governance Can Learn The NPT struck a bargain between the states that had the bomb and those that did not. Here is what that grand bargain teaches ASI governance.