Artificial Analysis has introduced a new agentic benchmark that evaluates large language models' ability to plan and execute tasks. Claude Fable and GLM 5.2 achieved top positions within their respective cohorts, demonstrating strong performance on this unsaturated benchmark.
New Agentic Benchmark Released
Anthropic releases Claude Opus 5 as a cost-effective alternative to Fable
Anthropic has released Claude Opus 5, a model designed to match the performance of Claude Fable 5 at approximately half the price per token while offering more permissive content classifiers. It serves as the new default model on Claude Max and the strongest option for Claude Pro users.
Kimi K3 ranks at same level as Opus thinking on Agent Arena
According to the Agent Arena leaderboard, Kimi K3 performs at a level comparable to Opus in non-vision tasks. While some users note that Opus may have an advantage in vision capabilities, the ranking indicates parity for other use cases.
SEA introduces anytime-valid certificates to self-evolving agents
The authors present SEA, an architecture that confines self-modification to a steering adapter and versioned harness around a frozen base model, admitting changes only through an anytime-valid gate that emits auditable certificates against a fixed error budget.
Anthropic releases Claude Opus 4.8 with adaptive reasoning and improved honesty
Anthropic has released Claude Opus 4.8, a model update that tops Artificial Analysis’s Intelligence Index and introduces adaptive thinking alongside parallel subagent workflows. The release also includes faster output modes and the ability to update system prompts mid-turn.
GLM-5.2 Outperforms GPT-5.5 in AA-Briefcase Evaluation
Artificial Analysis' new agentic knowledge work evaluation, AA-Briefcase, shows GLM-5.2 surpassing GPT-5.5 in performance. The benchmark assesses real-world task execution and reasoning capabilities in knowledge work scenarios.