HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench
HealthBench-Psych is a mental health subset extracted from OpenAI's HealthBench, yielding 610 conversations (12.2% of the 5,000-corpus) through an LLM-applied rubric and two rounds of blinded clinician validation HealthBench-Psych-Hard is introduced as a harder variant to better stress-test model performance on complex mental health scenarios Evaluation of 20 frontier and open models using a cross-vendor panel of three LLM judges found a statistically tied frontier cluster, measurable refusal be
Analysis
TL;DR
- HealthBench-Psych is a mental health subset extracted from OpenAI's HealthBench, yielding 610 conversations (12.2% of the 5,000-corpus) through an LLM-applied rubric and two rounds of blinded clinician validation
- HealthBench-Psych-Hard is introduced as a harder variant to better stress-test model performance on complex mental health scenarios
- Evaluation of 20 frontier and open models using a cross-vendor panel of three LLM judges found a statistically tied frontier cluster, measurable refusal behavior in two models, and near-identical rankings across judges (Kendall's tau ≥ 0.92)
- The paper addresses a critical gap: general health benchmarks cannot isolate mental health performance, while most existing mental health evaluations are bespoke academic benchmarks difficult to integrate into developer workflows
- All resources — subset, pipeline, model responses, grades, and analysis code — are released as a reusable open resource
Why It Matters
As millions of people increasingly turn to LLMs for psychological support, having a rigorous, clinically validated benchmark specifically for mental health evaluation is essential for both researchers and developers building AI-powered mental health tools. This work bridges the gap between academic benchmarks and production-grade evaluation by providing a reusable pipeline that can be integrated into standard developer workflows.
Technical Details
- Screening pipeline: Used a transparent LLM-applied rubric to screen 5,000 physician-rubricated conversations from HealthBench for mental health relevance, followed by two rounds of blinded clinician review with concealed known-exclude controls to validate the subset
- Subset composition: Yielded 610 conversations (12.2% of the corpus) forming HealthBench-Psych, with a harder variant (HealthBench-Psych-Hard) designed to challenge models on more complex clinical scenarios
- Evaluation framework: Assessed 20 frontier and open-source models using a cross-vendor panel of three LLM judges, ensuring diverse vendor representation to reduce bias
- Inter-judge reliability: Achieved near-identical rankings across all three judges (Kendall's tau ≥ 0.92), demonstrating strong consistency in the evaluation methodology
- Open release: The subset, screening pipeline, model responses, grades, and analysis code are all publicly released as a reusable resource
Industry Insight
- The measurable refusal behavior observed in two models highlights an ongoing tension in mental health AI: over-refusal can deny users support, while under-refusal risks harm — developers should carefully calibrate safety guardrails for mental health applications
- The statistically tied frontier cluster suggests that current state-of-the-art models have not yet achieved a meaningful performance gap in mental health evaluation, indicating the benchmark's discriminative power and the need for continued advancement
- The release of a reusable pipeline and open resources lowers the barrier for organizations to adopt rigorous mental health evaluation, potentially standardizing how AI mental health tools are assessed across the industry
Disclaimer: The above content is generated by AI and is for reference only.