Research Papers 论文研究 5h ago Updated 1h ago 更新于 1小时前 45

HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench HealthBench-Psych:OpenAI健康基准的心理健康子集

HealthBench-Psych is a mental health subset extracted from OpenAI's HealthBench, yielding 610 conversations (12.2% of the 5,000-corpus) through an LLM-applied rubric and two rounds of blinded clinician validation HealthBench-Psych-Hard is introduced as a harder variant to better stress-test model performance on complex mental health scenarios Evaluation of 20 frontier and open models using a cross-vendor panel of three LLM judges found a statistically tied frontier cluster, measurable refusal be 从OpenAI HealthBench的5,000个医生评分对话中筛选出610个心理健康相关对话(12.2%),构建HealthBench-Psych和HealthBench-Psych-Hard两个子集 采用透明LLM评分标准进行筛选,并通过两轮盲审临床医生审查验证,确保数据质量 评估20个前沿和开源模型,发现前沿模型集群统计上平局,两个模型存在可测量的拒绝行为 三个跨厂商LLM评委的排名高度一致(τ ≥ 0.92),验证了评估方法的可靠性 开源完整数据集、筛选管道、模型响应、评分和分析代码,便于社区复用

58
Hot 热度
72
Quality 质量
65
Impact 影响力

Analysis 深度分析

TL;DR

  • HealthBench-Psych is a mental health subset extracted from OpenAI's HealthBench, yielding 610 conversations (12.2% of the 5,000-corpus) through an LLM-applied rubric and two rounds of blinded clinician validation
  • HealthBench-Psych-Hard is introduced as a harder variant to better stress-test model performance on complex mental health scenarios
  • Evaluation of 20 frontier and open models using a cross-vendor panel of three LLM judges found a statistically tied frontier cluster, measurable refusal behavior in two models, and near-identical rankings across judges (Kendall's tau ≥ 0.92)
  • The paper addresses a critical gap: general health benchmarks cannot isolate mental health performance, while most existing mental health evaluations are bespoke academic benchmarks difficult to integrate into developer workflows
  • All resources — subset, pipeline, model responses, grades, and analysis code — are released as a reusable open resource

Why It Matters

As millions of people increasingly turn to LLMs for psychological support, having a rigorous, clinically validated benchmark specifically for mental health evaluation is essential for both researchers and developers building AI-powered mental health tools. This work bridges the gap between academic benchmarks and production-grade evaluation by providing a reusable pipeline that can be integrated into standard developer workflows.

Technical Details

  • Screening pipeline: Used a transparent LLM-applied rubric to screen 5,000 physician-rubricated conversations from HealthBench for mental health relevance, followed by two rounds of blinded clinician review with concealed known-exclude controls to validate the subset
  • Subset composition: Yielded 610 conversations (12.2% of the corpus) forming HealthBench-Psych, with a harder variant (HealthBench-Psych-Hard) designed to challenge models on more complex clinical scenarios
  • Evaluation framework: Assessed 20 frontier and open-source models using a cross-vendor panel of three LLM judges, ensuring diverse vendor representation to reduce bias
  • Inter-judge reliability: Achieved near-identical rankings across all three judges (Kendall's tau ≥ 0.92), demonstrating strong consistency in the evaluation methodology
  • Open release: The subset, screening pipeline, model responses, grades, and analysis code are all publicly released as a reusable resource

Industry Insight

  • The measurable refusal behavior observed in two models highlights an ongoing tension in mental health AI: over-refusal can deny users support, while under-refusal risks harm — developers should carefully calibrate safety guardrails for mental health applications
  • The statistically tied frontier cluster suggests that current state-of-the-art models have not yet achieved a meaningful performance gap in mental health evaluation, indicating the benchmark's discriminative power and the need for continued advancement
  • The release of a reusable pipeline and open resources lowers the barrier for organizations to adopt rigorous mental health evaluation, potentially standardizing how AI mental health tools are assessed across the industry

TL;DR

  • 从OpenAI HealthBench的5,000个医生评分对话中筛选出610个心理健康相关对话(12.2%),构建HealthBench-Psych和HealthBench-Psych-Hard两个子集
  • 采用透明LLM评分标准进行筛选,并通过两轮盲审临床医生审查验证,确保数据质量
  • 评估20个前沿和开源模型,发现前沿模型集群统计上平局,两个模型存在可测量的拒绝行为
  • 三个跨厂商LLM评委的排名高度一致(τ ≥ 0.92),验证了评估方法的可靠性
  • 开源完整数据集、筛选管道、模型响应、评分和分析代码,便于社区复用

为什么值得看

心理健康是公众健康关注焦点,数百万人转向LLM寻求心理支持,但现有评估基准难以隔离专科性能。该研究填补了心理健康领域LLM评估的空白,提供了可集成到开发者工作流的标准化工具。

技术解析

  • 数据筛选流程:使用透明LLM评分标准从HealthBench的5,000个对话中筛选心理健康相关内容,再通过两轮盲审临床医生审查验证,设置隐藏的控制样本确保筛选准确性
  • 子集规模:最终获得610个对话(占原始数据集12.2%),分为HealthBench-Psych和HealthBench-Psych-Hard两个版本
  • 评估方法:采用跨厂商的3个LLM评委组成评审小组,对20个前沿和开源模型进行交叉评估
  • 评估结果:前沿模型集群统计上无显著差异,两个模型表现出可测量的拒绝行为,评委间排名一致性高(Kendall's τ ≥ 0.92)
  • 开源资源:发布完整数据集、筛选管道代码、模型响应、评分结果和分析代码,支持社区复用和扩展

行业启示

  • 心理健康AI评估需要专科化基准,通用健康基准无法有效隔离特定领域的模型性能
  • LLM在心理健康场景的拒绝行为值得关注,可能影响用户获得支持的可用性
  • 跨厂商LLM评委评估方法具有高度一致性,为标准化评估提供了可行路径

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Benchmark 基准测试 Evaluation 评测 Healthcare AI 医疗AI LLM 大模型 Dataset 数据集