An AI Safety Institute is a government body set up to evaluate frontier AI systems for risks to national security and public safety, and to build the state's own expertise on frontier AI. The first wave clustered around the 2023 international AI summit. Institutes now share methods and findings through an international network that Seoul and follow-on meetings pushed forward.
Several major capitals now run an institute or an equivalent. Most of them still cannot order a dangerous model held back.
A note on names: the two founding institutes have been renamed. The UK body became the AI Security Institute in February 2025. The US institute is now the Center for AI Standards and Innovation (CAISI) at NIST. "AI safety institute" remains the standard generic term, and this guide uses it throughout.
For years the only organizations that deeply understood frontier models were the companies building them. An institute is an attempt to put independent technical capability inside government, so public authorities can assess systems rather than accept a developer's briefing as the last word.
What they actually do
Their work clusters into a few areas.
- Testing frontier models, sometimes before release, for dangerous capabilities in domains such as cyber, biology, and autonomy, using the methods behind capability evaluations.
- Developing the science of evaluation itself, because measuring these risks well is still an open research problem, not a settled procedure.
- Advising government, so policy is informed by people who have actually examined the systems.
- Coordinating internationally, so a model tested in one country need not be rebuilt from scratch everywhere, and so standards begin to converge.
This is real institutional progress. Building state capacity to understand frontier AI is a precondition for governing it. You cannot regulate what you cannot evaluate, and until recently governments largely could not.
The power they mostly lack
Most AI Safety Institutes can test, advise, and publish. Few can compel. Access to models often depends on lab cooperation. Findings usually inform rather than bind. In most cases an institute cannot order a dangerous model withheld. These bodies can see a risk and recommend a response. They rarely get to require one.
That gap between assessment and authority is the structural limit. An institute that discovers a serious hazard and can only advise is only as effective as the government's will to act against commercial and competitive pressure to keep going. Much of the current architecture is a warning system. A warning system is not a safeguard until someone with power answers the alarm.
The folk objection is that technical capacity is enough, and politics will catch up. Capacity without authority is how governments learn the bad news late and act later still. Testing that cannot force a stop is preparation for governance, not governance itself.
What they could become
The lasting value of AI Safety Institutes is what they assemble for binding regimes later: independent technical capacity, shared standards, and an international channel between governments. In that sense they are the early form of the kind of body the Foundation argues for. Our piece on an international monitoring agency describes where this could lead. The IPCC model shows how shared scientific assessment can underwrite international policy.
What has to change is authority. Assessment must connect to enforcement, whether through domestic law that makes an institute's sign-off a condition of deployment, or through an international framework that gives verified findings real consequences. Give these institutes power that matches their expertise, and much of the scaffolding of serious governance is already partly built. That transition is the subject of our plan: stop superintelligence under law, with public evaluation capacity as infrastructure, not as a substitute for prohibition and verification.