Qwen 3.8 Max achieves a score of 1588 on the Debate Benchmark, improving upon Qwen 3.7 Max's score of 1462. This benchmark evaluates how well large language models hold an argument under adversarial, multi-turn opposition across various topics.
- The Debate Benchmark measures broad knowledge, accurate fact usage under pressure, sharp rebuttal, and coherence over several rounds.
- Each matchup is debated twice with sides swapped to cancel side bias, and a three-model judge panel decides the winner and margin.
- Qwen 3.8 Max's average cost per debate increased by 45% compared to its predecessor.