OpenAI has temporarily slowed its model development pace, including a two-week pause on reinforcement learning (RL) training for upcoming models, after preliminary evidence indicated that its new model, Astra, may meet the Critical cybersecurity capability threshold. This decision follows the OpenAI-Hugging Face incident and aims to strengthen monitoring, alignment, and containment safeguards as AI systems become more capable.
- Workloads involving Astra or cyber models now require the strictest security safeguards, including stronger workload isolation (sandboxes) and network controls to prevent unauthorized internet access.
- A new multistage monitoring system uses activation classifiers to detect concerning behavior, aiming to issue alerts within 30 minutes for potential critical security boundary violations.
- Core alignment techniques are being applied across more training stages for the most capable models to better detect unsafe behavior, reward hacking, and deception.
- Monitoring overhead is estimated at roughly 20% of the inference compute being monitored, with safety teams paged for highest-priority alerts that cannot be resolved as false positives within 30 minutes.
OpenAI emphasizes that these changes are necessary because the risks associated with developing increasingly capable AI systems are growing, and their security standards must stay ahead of those risks to ensure alignment and safety.