Researchers present GPQA, a challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. The benchmark is designed to be extremely difficult; PhD-level experts achieve only 65% accuracy, while skilled non-experts reach 34% even with unrestricted web access. State-of-the-art AI systems also struggle, with the strongest GPT-4 baseline achieving just 39% accuracy. This difficulty for both humans and frontier models aims to enable realistic scalable oversight experiments for supervising AI outputs in scientific knowledge development.