OpenAI has introduced MentalHealthBench, an open benchmark designed to evaluate how AI systems respond in realistic mental health conversations. Developed with over 80 licensed mental health experts from 22 countries, the benchmark assesses model capabilities across key behaviors such as safety, context-seeking, and preserving user agency.
- The dataset includes synthetic conversations covering non-acute, high-acuity, and emergency scenarios for adults, teens, caregivers, and clinicians.
- Evaluation criteria were established by experts using a weighted rubric system, with responses graded by GPT-5.6 Sol.
- A separate analysis compared expert guidance with user preferences, revealing that users value practical next steps and tone more than experts do.
- The benchmark allows for multifaceted measurement across ten dimensions of model behavior to identify specific areas for improvement.
MentalHealthBench provides a shared tool for researchers and developers to examine gaps in existing models and work toward higher standards for safe, useful AI support in mental health contexts.