The Qwen 3.8 27B model, running in FP8 with a 256K context window via vLLM, achieved a score of 72.9 on the Aider benchmark.

  • The score matches Gemini 2.5 Pro (scored 72.9) and exceeds Claude Opus 4 (72.0) and DeepSeek R1 (71.4).
  • The author notes that while the benchmark may be outdated, the model's performance on a MacBook is comparable to state-of-the-art models from over a year ago.
  • Actual performance is described as surpassing these models due to improvements in the testing harness.