The UK AI Security Institute (AISI) has made publicly reported evaluation methods and findings available through EvalEval's Evaluation Cards platform to improve benchmark reproducibility. This release includes verified results, context, and configuration information for five specific benchmarks: HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro, and Terminal-Bench 2.0.
- The data covers six frontier models: Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2, and GPT-5.4.
- Additional results from two cyber evaluations, Cyber CTFs and The Last Ones, are included using a partially overlapping set of models.
- The release accompanies AISI's paper on how inference compute shapes frontier LLM evaluation, highlighting that performance on Humanity's Last Exam changes with evaluation protocol and token count.
Openly releasing results with setup information allows researchers to examine studies closely, compare findings across the ecosystem, and understand how setup choices influence reported performance.