Researchers present GPQA, a challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. The benchmark is designed to be extremely difficult; PhD-level experts achieve only 65% accuracy, while skilled non-experts reach 34% even with unrestricted web access. State-of-the-art AI systems also struggle, with the strongest GPT-4 baseline achieving just 39% accuracy. This difficulty for both humans and frontier models aims to enable realistic scalable oversight experiments for supervising AI outputs in scientific knowledge development.
GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Benchmarks
| Benchmark | Model | Score |
|---|---|---|
| GPQA Diamond | GPT-4 based baseline | 39% |
Researcher seeks one independent annotator to resolve LLM agreement ambiguity
A researcher is requesting a single independent human annotator to label 100 Turkish narrative scenes in order to determine whether low inter-rater agreement stems from the interpretive nature of the task or from underspecified annotation definitions. The study found that four LLMs and a rule-based detector agreed with each other and the human reference at roughly chance level, with Cohen’s κ scores ranging from 0.000 to 0.185.
Study identifies FragileTokens that fail contextual copying despite passing isolation probes
A study characterizes "FragileTokens," vocabulary entries in open-weight language models that successfully copy when isolated but exhibit errors when embedded in surrounding text. The research highlights that literal identity preservation is not guaranteed by standard isolation tests, as tokens can be deleted, substituted, or truncated within sequences.
GPT-5 and GPT-5 Nano match experts in microbial oncogenesis research appraisal
A study demonstrates that GPT-5 and GPT-5 Nano achieve expert-level performance in extracting evidence and critically appraising research publications on microbial oncogenesis. Researchers benchmarked these models alongside Gemini 2.5 Pro and Gemini 2.5 Flash against domain experts using a dataset of 24 papers focused on MMTV-LV and breast cancer.
GPT-5 and GPT-5 Nano match experts in microbial oncogenesis evidence extraction
A study benchmarking large language models on systematic evidence synthesis for microbial oncogenesis found that GPT-5 and GPT-5 Nano perform indistinguishably from domain experts. Researchers evaluated Gemini 2.5 Pro, Gemini 2.5 Flash, GPT-5, and GPT-5 Nano on 24 research papers using a structured template of 77 items across multiple question types.
M$^3$R-Bench introduces evidence-grounded benchmark for multimodal metaphor understanding
Researchers introduce M$^3$R-Bench, a unified benchmark containing 1,000 image-text instances with human-verified annotations designed to evaluate evidence-grounded multimodal metaphor understanding. The benchmark provides joint annotations for metaphor occurrence, Target-Source mapping, sentiment, and stage-wise explanations based on Conceptual Metaphor Theory.