Artificial Analysis has introduced a new agentic benchmark that evaluates large language models' ability to plan and execute tasks. Claude Fable and GLM 5.2 achieved top positions within their respective cohorts, demonstrating strong performance on this unsaturated benchmark.
New Agentic Benchmark Released
Anthropic finds GLM-5.3 achieves full control flow hijacks in cyber red teaming
Anthropic's Frontier Red Team evaluated the Chinese model GLM-5.3 on 100 tasks from an internal Binary Exploitation benchmark and found it capable of developing full control flow hijacks in 4% of trials.
Claude Opus 5.5 leads SimpleBench at 88.4% and excels in explainer videos
Claude Opus 5.5 has shipped, leading SimpleBench with an 88.4% score and generating significant community attention for its ability to produce high-quality explainer videos.
TrueForge achieves same accuracy as Claude Managed Agents with up to 75% lower cost
The developers of TrueForge, an open-source model-neutral agent harness, benchmarked it against Claude Managed Agents using the DevRev Enterprise-Bench. The comparison revealed that TrueForge can match the solve rate of managed platforms while significantly reducing resource consumption.
Anthropic launches Claude Fable and Mythos 5.1 with new SOTA benchmarks
Anthropic has released Claude Fable 5.1 and Claude Mythos 5.1 as its new flagship models for coding and knowledge work, positioning them as the world's most advanced models for autonomous, multi-step tasks.
Z.ai fine-tunes GLM-5.3 for cybersecurity gains
Z.ai released GLM-5.3, a 753-billion parameter model that improves coding and agentic capabilities through fine-tuning rather than training from scratch. The update ties the open-weights leader Kimi K3 on Artificial Analysis’ intelligence index and achieves top scores on cybersecurity benchmarks.