Jason Wei and coauthors at Google and elsewhere catalogued tasks where performance sat near chance until model scale crossed a threshold, then rose sharply. Multi-step arithmetic, specialized-domain questions, instructions in languages barely present in training, chains of reasoning steps: none of these was designed in as a feature. They showed up as models grew. The paper called them emergent abilities.
An emergent ability, in that usage, is a capability a smaller model does not have and a larger one does, where the transition looks less like a smooth ramp and more like a switch. Below some scale the model cannot do the task. Past that scale it can. The ability was not readable from the trend of smaller systems alone.
A real debate worth flagging
Emergence is not settled science. Some researchers argue that a share of reported jumps is partly a measurement artifact: a harsh all-or-nothing metric makes steady underlying progress look sudden, while a smoother metric shows gradual improvement. That critique lands on specific cases and deserves a fair hearing.
It does not dissolve the practical problem. Whether a capability arrives as a true discontinuity or as a steep curve noticed only after it crosses a threshold, the governance situation is the same. We learn a model can do something new after it can, not before. For a policymaker deciding whether the next training run is safe to green-light, the distinction between genuine emergence and a sharp curve is academic. Either way the surprise sits on the far side of the decision.
Why unpredictability is the core issue
We cannot reliably predict, before training a larger model, the full list of things it will be able to do. Scaling laws forecast broad performance measures such as loss reasonably well. They do not name which specific abilities will appear, or when. Capability is more predictable in aggregate than in particulars.
For most abilities that is merely interesting. For safety-relevant ones it is alarming. The same unpredictability that yields a surprise talent for translation can yield a surprise talent for the tasks covered in dangerous capability evaluations: cyber assist, bioweapon help, manipulation at scale. A capability no one anticipated is a capability no one tested for and no one built safeguards against.
We can forecast roughly how good the next model will be on average metrics. We cannot forecast everything it will be able to do. That gap is where the risk sits.
What it means for how we scale
Emergence turns each large jump in scale into a step into partial darkness. You build a system whose complete capability profile you learn only after it exists. By then, if a dangerous capability came with it, the system already has it. That is a different risk from a known hazard you can measure and mitigate in advance.
The Foundation's case follows directly. Frontier scaling should not run ahead of the ability to understand what is being created. The burden should fall on demonstrating that a large new system is safe before it is trained and deployed, not on discovering its dangers afterward. When abilities can arrive unannounced, caution before the jump is the only caution that helps. That framework is in our plan. The shortening timelines that raise the stakes are covered in the AGI timeline.