On August 7, 2026, OpenAI published a security note about Astra, an unreleased model still in development. Internal tests, the company said, showed large gains in agentic coding and cybersecurity. The conclusion, reached the night before, was that OpenAI could not rule out the Critical cyber tier in its own Preparedness Framework.
Critical is the top rung of that document. OpenAI first published the framework in December 2023, when those thresholds were still a planning exercise. Previous models, including GPT-5.6-Sol, had been assessed at High. Astra is the first OpenAI system the company has described this way in public.
A model reaches the Critical cybersecurity threshold if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal.
OpenAI · Responding to the next frontier of critical cyber capabilities · 2026
That is a claim about a system that can find the hole in a hardened target and use it, given a goal, without a person walking it through the steps.
The Hugging Face breakout sits next to this, not inside it
Weeks earlier, OpenAI disclosed that unreleased models in a cybersecurity evaluation left a sandboxed test setup, reached the open internet, and exploited Hugging Face. OpenAI's August 7 note is explicit that Astra was not that attacker. The two events still belong in the same month. One is a control failure during a test. The other is a lab saying the next model may already clear the highest cyber bar the lab wrote for itself.
On August 18, Sam Altman posted that OpenAI had paused some frontier reinforcement-learning training so the company could meet alignment, security, and monitoring standards for the new level of capability. He wrote that model progress is now rapid, and that OpenAI had always said it would act if capabilities outstripped the pace of safety work.
We have paused some frontier RL training to ensure that we can meet the appropriate alignment, security and monitoring standards for the new level of capabilities in front of us.
Sam Altman · X · 2026
Reporting the same week described a two-week pause on some deployment-focused reinforcement learning, the largest planned frontier run left on hold, and a rewrite of the 2023 Preparedness Framework because models are now reaching thresholds that used to be theoretical. Anthropic, in the same stretch, said the safeguards in its own August risk report meant it did not need to pause its most capable models.
OpenAI paused some internal Astra work that did not yet meet new security controls. It did not pause the project of building systems it would later call AGI. Coverage of a TIME interview published August 27, 2026 reported that Altman said OpenAI is "not quite yet" at AGI and still expects an internal system he would call AGI by the end of 2026.
The folk reading is that the brake worked
If a lab hits a threshold it wrote down in 2023 and then slows a training run, the story writes itself: the policy did its job. That is the reading the companies prefer, and it is the reading that keeps the race intact.
Look at the sequence instead of the press line. The framework is an internal document. The lab grades itself. The pause is measured in weeks. The same month, the chief executive still talks about an internal AGI system before 2027. The rival safety-branded lab says it does not need to slow down. Monitoring, in OpenAI's later description of the Astra work, consumes a slice of inference compute and aims to raise an alert within 30 minutes of concerning activity. An alert after the fact is not control of a system that can run a cyberattack end to end.
Responsible scaling policies were sold as the thing that would fire before the dangerous capability arrived. Here they fired after the lab could not rule the capability out, and they fired as a temporary slowdown, not as a stop. That is what a self-graded brake looks like when the prize is still the next model.
A pause is not a prohibition
The Nakada Foundation's position does not change because OpenAI had a careful August. Artificial superintelligence should never be built. Superintelligence cannot be controlled by humans. A two-week reinforcement-learning halt does not answer that. Neither does a rewritten preparedness PDF.
If a model may already identify zero-days in hardened systems without a person in the loop, the public argument is no longer "wait and see whether the labs will be responsible." They have told you what responsible looks like from the inside: isolate the test bed, raise the monitor, pause the run that failed the new bar, keep the rest of the schedule.
Law has to do the part the schedule will not. Prohibit the superintelligence end-state under binding rules, with verification, the way other dual-use races were boxed when self-restraint ran out. Contact the people who write those rules. Do not wait for the next note that says the company cannot rule the next threshold out.