GLM-5.2, an open-weight AI model released by Z.ai, has set a new benchmark in coding and general agent performance. It outperforms models like Claude Fable 5 and Gemini, and matches or exceeds OpenAI's Opus 4.8 in max thinking mode, establishing itself as the first open model that feels right in coding harnesses as a general agent.
GLM-5.2 is the step change for open agents
Z.ai releases GLM-5.3 with improved coding and cybersecurity via post-training
Z.ai has released GLM-5.3, an update to its 743B parameter model that achieves performance gains through scaled post-training rather than base model retraining. The release is currently available via the Z.ai API, GLM Coding Plan, and ZCode, with public weights scheduled for release approximately two weeks after launch following safety evaluations.
xAI launches Build Mode; researchers urge AI slowdown pact
xAI has launched "Build Mode" for SuperGrok Heavy subscribers, enabling users to generate, edit, preview, and publish websites, apps, games, and dashboards directly from chat without setup. Simultaneously, over a thousand employees from frontier AI companies have signed a statement requesting that the US government support an international effort to develop tools for deliberately pacing the frontier of automated AI development.
ZCode: New Agentic Code Editor from the Makers of GLM
The creators of GLM have released ZCode, a new agentic code editor. The tool is available at zcode.z.ai.
Qwen3.6-27B with 3-Critic Harness Matches Frontier Quality
A user tested Qwen3.6-27B (8-bit) alongside GLM5.2 using a coding harness that employs three critics—code review, test review, and Playwright e2e—to validate output quality.
GLM-5.2 Breakout and Open-Model Progress Highlighted
Zhipu's GLM-5.2 emerged as the top open-weight model, praised for its frontier-adjacent performance in daily use, with improvements in coding tasks and reduced 1M-token inference cost via IndexShare. It outperformed other open models in agentic knowledge work benchmarks, reaching 1266 Elo in Artificial Analysis' AA-Briefcase test, though only 3% of tasks were fully satisfied by top models, indicating persistent challenges in real-world long-horizon agent performance.