The ICML 2026 Open Reproductions challenge, held from July 15 to August 2, 2026, engaged the community in reproducing accepted papers using AI agents and human oversight. The effort resulted in the largest open, claim-by-claim audit of a machine learning conference to date.
- Of the examined papers, 51% had at least one claim independently verified, while 23% had at least one claim falsified or contested.
- 496 papers had at least one falsified or contested claim, including 49 where all claims were falsified and 242 where teams reached opposite verdicts.
- Confirmed errors included a paging algorithm's robustness term growing incorrectly, a theorem failing after step 224, and code using the wrong KL divergence loss.
- The challenge highlighted that human-in-the-loop workflows were more reliable than pure agent execution for catching subtle errors and evaluating perceptual quality.
The organizers conclude that while agents can scale reproduction efforts, humans remain essential for steering agents, questioning assumptions, and performing irreducibly human evaluations.