GLM-5.2 is the first open-weights model to achieve 80% accuracy on Terminal-Bench and outperforms all other available open models. It also surpasses Gemini, positioning it as a frontier-level model at a significantly lower cost.
GLM-5.2 crosses 80% on Terminal-Bench
Benchmarks
| Benchmark | Model | Score |
|---|---|---|
| Terminal-Bench | GLM-5.2 | 80% |
Z.AI confirms Ox Alpha is a new GLM series model and will release weights
Z.AI confirmed on Wednesday that the Ox Alpha model is a new iteration of its GLM series. The company stated it will release the model's weights later that night, responding to queries by Bloomberg News.
Z.ai releases GLM-5.3 with extended post-training
Z.ai announced GLM-5.3, a new model that achieves frontier performance on agentic coding benchmarks by extending the post-training of its GLM-5.2 base model. The ~750B parameter model currently surpasses Kimi K3 and matches or exceeds Claude Fable 5 and GPT-5.6-Sol on various benchmarks.
Kimi K3, DeepSeek V4 Pro, and GLM-5.2 compared on benchmarks, license, and serving cost
A comparison of three open-weight sparse Mixture-of-Experts models—Moonshot AI's Kimi K3, DeepSeek V4 Pro, and Zhipu AI's GLM-5.2—evaluates their capabilities, licensing terms, and serving costs for long-horizon coding and agent workloads.
Human Evaluation Shows GLM-5.2 Competes with Top Models
A human evaluation on Design Arena's leaderboard reveals GLM-5.2 performs nearly as well as Fable 5 in game development tasks, placing just one step below it. The model, based on open weights and MIT licensing, is assessed as equivalent in capability to the best available Claude models, suggesting that standardized benchmarks may no longer accurately reflect real-world performance.
GLM-5.2 Is the New Best Open Model
GLM-5.2 achieves benchmark scores near frontier levels, matching Opus 4.7 in text-only tasks and ranking among the top open models on multiple tests. It is the strongest open model currently available, outperforming predecessors and rivals like GPT-5.5 and Fable, though it falls short on specialized benchmarks like anti-sycophancy and has limited vision capabilities.