Researchers introduce BenchMIRT, a method for auditing large language model benchmarks at the individual prompt level using multidimensional Item Response Theory. By analyzing performance across 100 models and 34,000 questions, the tool disentangles mixed capabilities like safety and general reasoning that often obscure benchmark scores.
- BenchMIRT independently recovered safety and general reasoning as dominant dimensions without prior labeling.
- It revealed that benchmarks like BBQ and WMDP correlate more strongly with general reasoning than their intended safety focus suggests.
- The method identifies the most informative questions, showing that retaining just 10% of items can preserve nearly the same capability assessment.
- BenchMIRT predicts held-out question correctness with 79% accuracy, outperforming simpler baseline approaches.
This approach helps researchers understand what benchmarks actually measure and design more efficient, focused evaluations by removing less informative questions.