Researchers introduce MMLU-Pro, an enhanced version of the Massive Multitask Language Understanding benchmark designed to address performance saturation and instability in current large language models. The new dataset expands the choice set from four to ten options, integrates challenging college-level reasoning questions, and eliminates noisy items through expert review.

  • Accuracy drops by 16% to 33% compared to MMLU, with top model GPT-4o scoring only 72.6%.
  • Model score sensitivity to prompt variations decreases from 4-5% in MMLU to just 2% in MMLU-Pro across 24 styles.
  • Chain of Thought reasoning significantly boosts performance on MMLU-Pro, improving GPT-4o by 19%, whereas it hurts results on the original benchmark.
  • The benchmark spans 14 domains with over 12,000 questions, providing greater discriminative power between models like GPT-4o and GPT-4-Turbo.

MMLU-Pro serves as a more reliable and difficult metric for tracking progress in expert-level language understanding and reasoning capabilities.