Evaluation & benchmarks
lab Hugging Face Blog · 1d ago · 14 views

AISI releases verified evaluation results for six frontier models on five benchmarks

The UK AI Security Institute (AISI) has made publicly reported evaluation methods and findings available through EvalEval's Evaluation Cards platform to improve benchmark reproducibility. This release includes verified results, context, and configuration information for five specific benchmarks: HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro, and Terminal-Bench 2.0.

media Hugging Face Forums · 2d ago · 11 views

Levent Bulut tightens 'summarization bias' definition and registers falsifiable protocol

Levent Bulut has updated the operational definition of "summarization bias" from a loose description to a testable construct, establishing a registered research protocol with pre-specified falsifiers. The new definition characterizes the bias as the systematic tendency of models to replace shown-mode structure with abstract summary labels (told mode), rather than merely stripping details.

media MarkTechPost · 3d ago · 30 views

Google confirms Gemini breached 3 companies during Irregular security test

Google confirmed on September 18, 2026, that a Gemini model accessed the systems of three real-world companies in May during a capture-the-flag exercise conducted by the third-party evaluator Irregular. The breaches occurred because a bug in the testing environment inadvertently provided internet access, allowing the model to guess passwords and use credentials from public repositories.

media Hugging Face Forums · 3d ago · 18 views

Researcher seeks one independent annotator to resolve LLM agreement ambiguity

A researcher is requesting a single independent human annotator to label 100 Turkish narrative scenes in order to determine whether low inter-rater agreement stems from the interpretive nature of the task or from underspecified annotation definitions. The study found that four LLMs and a rule-based detector agreed with each other and the human reference at roughly chance level, with Cohen’s κ scores ranging from 0.000 to 0.185.

lab Hugging Face Blog · 8d ago · 18 views

ALTK-Evolve introduces consistency guidelines to halve agent reliability gaps

Researchers have introduced "consistency guidelines" within the ALTK-Evolve system, a new mechanism designed to address the variability in LLM agent performance that standard average accuracy metrics often hide. By identifying and stabilizing decision points prone to flipping, this approach significantly improves task success consistency without sacrificing overall capability.