Anthropic has launched the Claude Opus 5 model, triggering scrutiny of its performance on coding and general capability metrics alongside renewed debate about frontier model evaluation.

  • Epoch AI reported Claude Opus 5 achieves an ECI of 159, slightly below Fable 5's 161, while matching Fable 5 with a SWE-ECI of 161 on software engineering benchmarks.
  • Users criticized the ECI score as understating practical improvements, noting it is only one point better than Opus 4.8 despite perceived qualitative gains.
  • An evaluation irregularity was noted where Opus 5 scored better on FrontierCode at medium effort than at higher effort, suggesting non-monotonic gains from extra inference-time compute.
  • Microsoft CTO Kevin Scott reported a head-to-head win against Fable using "best-of-n" sampling, while Nous Portal announced access to the model with a 20% discount.

The launch highlights the tension between aggregate benchmark scores and user-reported performance in agentic tasks like browser automation.