OpenAI says its next major model may be dangerous enough to write its own cyberweapons, and it's pulling back until the safeguards catch up.
“Our latest internal evaluations of Astra, one of our upcoming models, over the past few days indicate significant advancements in agentic coding and cybersecurity,” OpenAI said. “These results, in addition to expert assessments, have led us to conclude last night that we cannot rule out critical cyber capabilities under our Preparedness Framework (opens in a new window).”
The framework is OpenAI's rulebook for risky models, first published in December 2023. "Critical" is its top rung. A model hits it if it can find and build working zero-day exploits (previously unknown holes a vendor hasn't patched) across hardened systems without a human in the loop, or if it can plan and run a full attack on a tough target from nothing but a high-level goal. Earlier models, including GPT-5.6-Sol, topped out at the lower "High" tier.
The pattern is already realOpenAI's caution reads differently once you line it up against what has been happening across the last few weeks. This isn't a future worry. Frontier models have already broken out of their test cages and gone after live targets.
The UK's AI Security Institute found the behavior wasn't a one-off, either. During testing of Anthropic's Mythos 5 and OpenAI's GPT-5.6-Sol, it logged 10 instances in 122 where the models took unsanctioned action on the live internet, one of them trying to slip malicious code into an open-source project.
OpenAI's response to Astra is to lock the door before the model is ready. It's pausing internal Astra work that lacks the new controls, isolating test environments, restricting network and tool access, protecting model weights, and monitoring risky actions across the board.


















