Researchers present GPQA, a challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. The benchmark is designed to be extremely difficult; PhD-level experts achieve only 65% accuracy, while skilled non-experts reach 34% even with unrestricted web access. State-of-the-art AI systems also struggle, with the strongest GPT-4 baseline achieving just 39% accuracy. This difficulty for both humans and frontier models aims to enable realistic scalable oversight experiments for supervising AI outputs in scientific knowledge development.
GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Benchmarks
| Benchmark | Model | Score |
|---|---|---|
| GPQA Diamond | GPT-4 based baseline | 39% |
M$^3$R-Bench introduces evidence-grounded benchmark for multimodal metaphor understanding
Researchers introduce M$^3$R-Bench, a unified benchmark containing 1,000 image-text instances with human-verified annotations designed to evaluate evidence-grounded multimodal metaphor understanding. The benchmark provides joint annotations for metaphor occurrence, Target-Source mapping, sentiment, and stage-wise explanations based on Conceptual Metaphor Theory.
Zing framework improves LLM social intelligence via SoMBench benchmark and Actio grounding
The report introduces Zing, an integrated framework designed to enhance the social intelligence of large language models by measuring, internalizing, and grounding social capabilities. The authors present SoMBench, a psychology-grounded benchmark spanning 17 secondary dimensions and 3,481 expert-verified instances, which reveals that current LLMs have significant room for improvement with no model reaching near-ceiling performance.
SRRM benchmark reveals Qwen2.5-1.5B outperforms 3B in long-context memory
A researcher is seeking an arXiv cs.CL endorsement for a paper introducing Simple Retrieval-Reconstruction Memory (SRRM), a lightweight benchmark designed to evaluate long-context memory in Small Language Models.
LLM fine-tuning with traits improves cross-rubric essay scoring
Researchers address the challenge of automated essay scoring when educators revise or introduce new rubrics, a scenario known as cross-rubric generalization. They propose a fine-tuning framework for Large Language Models that utilizes rubric-agnostic intermediate representations called "traits" alongside target-essay supervision during training.
Live Gurbani Tracking: Benchmark and Reference System for Sikh Kirtan Captioning
The authors present a benchmark and reference system for the live captioning of Sikh Kirtan, which involves the continuous sung recitation of verses from the Sri Guru Granth Sahib Ji. Unlike open-vocabulary transcription, this task requires exact word-for-word matches from the canonical scripture to avoid religiously inappropriate misspellings.