Before a new drug ships in the United States, someone has to show safety and efficacy under rules that can stop a launch. Most software still ships first and apologizes later. The FDA pattern is the question AI policy keeps dodging: prove it before the damage is irreversible.

Pharmaceuticals are the great exception. A drug company cannot release a medicine and wait for the bodies. It has to demonstrate safety and efficacy to the Food and Drug Administration through staged clinical trials before it may sell the product at all. The burden of proof sits with the developer, and it sits before deployment. The FDA model for AI asks whether frontier AI, which shares the features that made medicine an exception, should be governed the same way.

Why the analogy is attractive

The fit is closer than it first appears, because the reasons drugs get pre-market approval are reasons that apply to frontier AI.

  • The burden of proof is reversed. The developer must show the product is safe, rather than a regulator or the public having to show it is dangerous after release. That inversion is exactly what the Foundation argues for elsewhere, in safety cases and scaling policies.
  • Approval comes before deployment. The check happens while the system is still contained, not after it is loose in the world, which is the only point at which prevention beats reaction.
  • Staged testing scales with risk. Clinical trials proceed in phases, each gate requiring evidence before the next. That maps naturally onto capability thresholds for AI.

For a technology where some failures may be severe and hard to reverse, a regime that refuses to let the highest-risk systems out until safety is affirmatively demonstrated is a serious and attractive proposition.

Where the model strains

The analogy is a guide, not a template, and the differences are instructive.

A drug is a fixed molecule doing a specific thing in a body, studied through a mature science with quantified risks. A frontier model is general-purpose, used for open-ended tasks nobody fully enumerated, with capabilities that can emerge unpredictably and behavior that can shift after deployment through fine-tuning or new tools. You can define what it means for a blood-pressure drug to be safe. Defining what it means for a general reasoning system to be safe is the unsolved problem at the center of the field, and the science an FDA-style reviewer would need often does not yet exist.

There is also the border problem. The FDA governs a national market with hard edges. A dangerous model can be trained anywhere and copied everywhere, so national pre-approval alone leaves the gap that only international coordination can close, backed by the physical chokepoints of compute governance.

The FDA model gives us the right principle, prove safety before release, and reminds us how much of the underlying safety science we still lack.

What to take from it

The FDA model is valuable less as a blueprint to copy than as a demonstration that pre-market approval is normal, workable, and accepted for products where the downside is too serious to handle by recall. Society already agrees, in the case of medicine, that some things must be proven safe before they reach us. Extending that settled principle to the frontier AI systems whose failures could be gravest is the application of an existing norm to a new technology that plainly qualifies, delivered through the mix of domestic requirement and international framework set out in our plan, not a radical demand.