Anthropic has paused internal activities involving its upcoming model, Astra, because recent evaluations indicate it may possess critical cybersecurity capabilities under the company's Preparedness Framework. The firm cannot rule out that the model can identify and develop functional zero-day exploits in hardened real-world systems without human intervention.
- Astra reached a threshold where critical cyber capabilities cannot be ruled out based on internal evaluations and expert assessments.
- Internal activities involving Astra are paused until they meet strengthened security control requirements.
- Stricter controls include isolated testing environments, restricted network access, enhanced weight protections, and universal monitoring for risky actions.
- The company is scaling up robustness testing of safeguards and will collaborate with government agencies and AI safety organizations.
Anthropic states that advanced cyber-capable models should help defenders identify vulnerabilities before attackers do, aiming to ensure frontier capabilities are deployed responsibly.