OpenAI introduced GPT-6 Sol and Luna as faster, more affordable counterparts to GPT-6 Astra, bringing advances in coding, factuality, computer use, and professional tasks to lower-cost models.

  • Anthropic released Claude Opus 5.5, which matches Claude Fable 5.1 on most work while costing 40% less to run than Opus 5.
  • SWE-Bench Pro V2 releases with 642 tasks from 11 repositories, correcting previous task errors and optimizing evaluation processes.
  • Google introduced RRSI, a method that regularizes how AI agent harnesses recursively improve themselves to reduce benchmark overfitting.

These updates provide users with more cost-effective model options and rigorous new benchmarks for evaluating AI performance.