A study auditing six widely used physics benchmarks finds that reported low scores for frontier language models are largely due to benchmarking errors rather than model deficiencies. Experts reviewed problem statements, reference solutions, and model responses, distinguishing genuine errors from grader mistakes, incorrect references, and ambiguous questions.
- GPT-5.6-Sol's mean@4 on HLE-Physics rose from 47.3% to 78.7% after correction.
- Its score on CMT-Benchmark increased from 61.0% to 87.2% following expert review.
- The model achieved a corrected pass@4 of 94.4% on the 54 retained CritPt challenges.
- Scores on UGPhysics, PRISM-Physics, and PHYBench also rose substantially after auditing.
The findings indicate that current benchmarks substantially understate frontier models' ability to solve well-posed physics problems, highlighting the need for more demanding, expert-validated evaluations.