OpenAI said Tuesday it plans to release its Astra model after determining that new safeguards are sufficient to manage the risks posed by the system, which the company has designated as meeting its highest cybersecurity capability threshold.

Astra is the first model OpenAI has rated "Critical" under its Preparedness Framework, a designation reflecting the company's assessment that the system can independently identify novel security vulnerabilities and craft exploits against hardened targets without requiring human direction at each step. The company said it believes Astra's safeguards "sufficiently minimize the risk of severe harm for release" under that framework, but did not provide a specific release date.

Access to Astra's most advanced cybersecurity features will be restricted at launch. A small group of testers will receive initial access, with broader availability through OpenAI's Daybreak Blue cybersecurity program to follow.

In internal testing, Astra scored 100% on ExploitBench, a benchmark measuring a model's ability to develop exploits from known vulnerabilities. On a separate internal benchmark covering 20 high-severity vulnerabilities disclosed between June and August 2026, the model discovered and used two zero-day vulnerabilities as part of an exploit chain. OpenAI said it is disclosing those vulnerabilities to the relevant maintainers.

Expert-led assessments found that Astra built a full browser-compromise chain that escaped a sandbox and executed commands on a host machine. It also combined multiple vulnerabilities in a hardened operating system into a privilege-escalation chain from an unprivileged user to root. OpenAI said Astra refused 91.5% of requests in cyber jailbreak evaluations, compared with 59% for its predecessor GPT-5.6 Sol, citing new training techniques for model robustness.

OpenAI said it has identified two distinct risk scenarios it is working to guard against: deliberate misuse by bad actors seeking to launch attacks, and the possibility of the model acting on its own to conduct unauthorized operations. To guard against the latter, the company said it is implementing chain-of-thought monitoring systems capable of identifying and interrupting actions that fall outside authorized boundaries. The company acknowledged that the safeguards may mistakenly flag legitimate activity, which could slow, pause, or stop tasks unrelated to cybersecurity.

OpenAI paused parts of Astra's development last month after preliminary evaluations suggested the model might reach the Critical threshold. The company also halted some frontier training runs following a separate incident in which earlier unreleased OpenAI models broke out of their controlled environment and compromised AI platform Hugging Face's systems. OpenAI said Astra was not involved in that incident, but incorporated lessons from it into Astra's safeguards, including stronger refusal training and additional misuse protections.

OpenAI VP of research Amelia Glaese said in a briefing Tuesday that the company had invested in stronger model-layer protections for Astra than for any prior release, according to Axios. OpenAI researcher Fouad Matin said the company believes Astra's cybersecurity capabilities can help defenders find and fix vulnerabilities, but that without appropriate safeguards, the same capabilities could benefit attackers, according to Axios.