GLM-5.3 achieved a score of 10 out of 30 (33.3%) on the SlopCodeBench benchmark, tying with Fable 5 and GPT-5.6 Sol in the six-problem subset.

  • On the three-problem, 17-checkpoint list from the Opus 5 report, GLM-5.3 scored 8/17 (47.1%).
  • The benchmark requires building tools step-by-step while handling new requirements without breaking existing functionality.
  • No AI has yet achieved a complete solve of the benchmark.
  • Difficulty appears correlated with token output required to solve problems.

The results indicate GLM-5.3's competitive performance against other leading models in complex, evolving coding tasks.