Muse-Ltd has submitted UncertaintyGym to the Hugging Face Evaluation Hub, a new benchmark designed to evaluate large language models' meta-cognitive calibration and their ability to express uncertainty rather than hallucinate.
The benchmark categorizes queries into four distinct types: solvable factual questions, under-specified queries requiring disambiguation, false premise questions needing rejection, and inherently unknowable topics. It utilizes the Meta Cognitive Calibration Score (MCS) and Unanswerable Hallucination Rate as primary metrics, with initial baselines showing LiquidAI's LFM2.5-2.6B model achieving a 45.0% MCS.
This submission aims to provide a standardized method for assessing how well models handle missing context or impossible premises by explicitly declaring unknowability.