Terminal Bench 4.0 has been released, updating the benchmark to keep pace with new model releases and combat benchmark saturation.

  • The update includes performance data showing GLM-5.3 is at the same level as Fable 5, accounting for margin of error.
  • The announcement emphasizes rapid iteration on TerminalBench to maintain relevance against emerging models.
  • Community discussion highlights the high computational cost of large benchmarks, noting they require 5-10B tokens per run.