OpenAI has designated its new model, Astra, as meeting the "Critical" cybersecurity capability threshold under its Preparedness Framework, meaning it can identify and exploit security flaws in hardened systems without human guidance. To mitigate risks, the company delayed parts of Astra's development to implement stronger safety measures, including improved alignment training and monitoring systems.

  • Astra achieved a perfect score on ExploitBench and discovered two zero-day vulnerabilities during internal testing.
  • The model refuses 91.5% of cyber jailbreak requests, compared to 59% for GPT-5.6 Sol.
  • Astra demonstrated the ability to build browser-compromise chains and local privilege-escalation exploits in expert-led assessments.
  • Access to advanced capabilities will initially be limited to alpha testers via Daybreak Blue.

OpenAI aims to balance the benefits of Astra's defensive utility with rigorous controls to prevent misuse or unauthorized actions by the model itself.