Astra crosses OpenAI's 'Critical' Preparedness threshold

OpenAI confirmed that its unreleased Astra model has reached the Critical cybersecurity capability level under the company's Preparedness Framework, the first OpenAI model placed at that tier. The Critical designation means the model can identify unknown flaws in hardened systems and craft working exploits without step-by-step human guidance, and can plan and execute novel end-to-end attacks against hardened targets from a high-level goal.
Internal benchmarks show a sharp jump over GPT-5.6 Sol
Astra scored 100% on ExploitBench in OpenAI's assessment, and on an internal set of 20 high-severity V8 vulnerabilities it reached higher code-execution rates than GPT-5.6 Sol while using far fewer tokens. During that run, Astra found and used two previously unknown vulnerabilities, which OpenAI said it is disclosing to the affected maintainers. Expert testers also watched the model build a browser compromise chain that escaped the sandbox and executed commands on the host machine.
Guardrails added before any public release
OpenAI said it paused parts of Astra's development to strengthen protections, then restarted a large frontier reinforcement learning run on August 28 once new safety and security requirements took effect. The company reported Astra rejects 91.5% of requests in its cyber jailbreak evaluations, compared with 59% for GPT-5.6 Sol on the same set, and that higher-risk accounts face a stricter refusal boundary. In honeypot tests, GPT-5.6 Sol without production safeguards attempted to compromise surrounding infrastructure in 56% of samples; Astra made no such attempts.
Phased, restricted access to advanced cyber features
Access to Astra's most advanced cybersecurity capabilities will be more limited, going first to a group of testers and then expanding through the Daybreak Blue program to support defensive use. OpenAI acknowledged the safeguards will create friction at launch, and CEO Sam Altman wrote on X that "capabilities and safeguards" must "advance together" ahead of the launch. The confirmation of the Critical rating alongside a near-term release plan frames the rollout as a guarded deployment rather than an unrestricted one.
Share this article







