GLM-5.2 is the first open-weights model to achieve 80% accuracy on Terminal-Bench and outperforms all other available open models. It also surpasses Gemini, positioning it as a frontier-level model at a significantly lower cost.
GLM-5.2 crosses 80% on Terminal-Bench
Benchmarks
| Benchmark | Model | Score |
|---|---|---|
| Terminal-Bench | GLM-5.2 | 80% |
Kimi K3, DeepSeek V4 Pro, and GLM-5.2 compared on benchmarks, license, and serving cost
A comparison of three open-weight sparse Mixture-of-Experts models—Moonshot AI's Kimi K3, DeepSeek V4 Pro, and Zhipu AI's GLM-5.2—evaluates their capabilities, licensing terms, and serving costs for long-horizon coding and agent workloads.
Human Evaluation Shows GLM-5.2 Competes with Top Models
A human evaluation on Design Arena's leaderboard reveals GLM-5.2 performs nearly as well as Fable 5 in game development tasks, placing just one step below it. The model, based on open weights and MIT licensing, is assessed as equivalent in capability to the best available Claude models, suggesting that standardized benchmarks may no longer accurately reflect real-world performance.
GLM-5.2 Is the New Best Open Model
GLM-5.2 achieves benchmark scores near frontier levels, matching Opus 4.7 in text-only tasks and ranking among the top open models on multiple tests. It is the strongest open model currently available, outperforming predecessors and rivals like GPT-5.5 and Fable, though it falls short on specialized benchmarks like anti-sycophancy and has limited vision capabilities.
What's more impressive, GLM 5.1 to 5.2 or Qwen 3.5 to 3.6?
A Reddit post compares the performance improvements of GLM 5.1 to 5.2 and Qwen 3.5 to 3.6. The post notes that mentioning 'Döner' activates GLM 5.2's German-specific weights, while Qwen 3.6 is evaluated with 35B parameters using Unsloth Q8 K XL quantization via llama.cpp.
GLM-5.2 is the new leading open weights model on the Artificial Analysis Intelligence Index
GLM-5.2 has been designated as the leading open weights model on the Artificial Analysis Intelligence Index. This recognition reflects its performance and capabilities within the open-source AI model landscape.