Research Papers 论文研究 3h ago Updated 47m ago 更新于 47分钟前 48

Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety 偏好而非更安全:成对偏好是临床安全的糟糕代理

Clinician pairwise preferences are a poor proxy for clinical safety in LLM evaluation, with high-preference models still showing substantial clinically meaningful failures Safety-critical failures (rated ≤ -1 on Harmlessness and Accuracy) are unevenly distributed across medical specialties, creating invisible domain-specific "no-go zones" Surface-level features explain slightly more preference variation than actual safety-critical rubric differences, and many preference votes carry no positive s 临床医生成对偏好是LLM临床安全性能的差代理指标,高排名模型仍可能存在显著的临床失败 临床失败在专科间分布不均,形成聚合排名和单数字排行榜中不可见的领域特定"禁区" 表面特征比安全关键特征更能解释偏好变化,大量偏好投票不含积极安全信号 提出结合成对偏好与评分反馈的临床调整排名方法,优于纯Bradley-Terry强度排序 研究基于MOOVE平台26,804个成对判断、13个LLM、736+临床医生、28+国家的真实数据

62
Hot 热度
78
Quality 质量
70
Impact 影响力

Analysis 深度分析

TL;DR

  • Clinician pairwise preferences are a poor proxy for clinical safety in LLM evaluation, with high-preference models still showing substantial clinically meaningful failures
  • Safety-critical failures (rated ≤ -1 on Harmlessness and Accuracy) are unevenly distributed across medical specialties, creating invisible domain-specific "no-go zones"
  • Surface-level features explain slightly more preference variation than actual safety-critical rubric differences, and many preference votes carry no positive safety signal
  • The authors introduce a clinically adjusted preference ranking that combines pairwise preference with rubric-derived feedback for more safety-aware model ordering
  • The study analyzed 26,804 pairwise judgments from 736+ clinicians across 28+ countries evaluating 13 LLMs on the MOOVE platform

Why It Matters

This research directly challenges a widely adopted evaluation methodology in clinical AI, where pairwise preference ranking has become a standard proxy for safety assessment. For AI practitioners building or deploying clinical LLMs, relying solely on preference-based leaderboards may create a false sense of security while dangerous domain-specific failures remain undetected. The findings have immediate implications for how clinical AI systems are evaluated, benchmarked, and regulated.

Technical Details

  • Dataset: 26,804 blinded pairwise judgments from 736+ clinicians across 28+ countries, comparing outputs from 13 LLMs on the MOOVE (Massive Open Online Validation and Evaluation) platform
  • Evaluation scale: Clinicians assign discrete scores on a [-2, +2] scale, where negative values indicate clinically unsafe or misleading content
  • Key dimensions analyzed: Harmlessness and Accuracy, with feature decomposition examining prompt length, refusal/escalation behavior, and surface-level vs. safety-critical feature contributions
  • Methodology: Bradley-Terry strength modeling for pairwise preference ranking, combined with multi-criterion rubric ratings for the clinically adjusted preference ranking
  • Novel contribution: A clinically adjusted preference ranking method that integrates pairwise preference with rubric-derived safety feedback, outperforming raw Bradley-Terry strength alone

Industry Insight

  • Evaluation frameworks for clinical LLMs must decouple preference ranking from safety assessment; single-number leaderboards are insufficient and potentially dangerous for clinical deployment decisions
  • Organizations should adopt multi-dimensional safety reporting that surfaces domain-specific failure rates rather than relying on aggregate preference scores, particularly for high-stakes medical specialties
  • The clinically adjusted preference ranking methodology offers a practical template for other high-stakes domains (legal, financial, autonomous systems) where preference-based evaluation may similarly mask critical safety failures

TL;DR

  • 临床医生成对偏好是LLM临床安全性能的差代理指标,高排名模型仍可能存在显著的临床失败
  • 临床失败在专科间分布不均,形成聚合排名和单数字排行榜中不可见的领域特定"禁区"
  • 表面特征比安全关键特征更能解释偏好变化,大量偏好投票不含积极安全信号
  • 提出结合成对偏好与评分反馈的临床调整排名方法,优于纯Bradley-Terry强度排序
  • 研究基于MOOVE平台26,804个成对判断、13个LLM、736+临床医生、28+国家的真实数据

为什么值得看

这项研究直接挑战了当前LLM评估中广泛使用的成对偏好方法在临床安全场景下的有效性,揭示了"偏好≠安全"的关键问题。对医疗AI从业者和评估框架设计者而言,研究提供了实证依据,推动建立更可靠的安全评估实践。

技术解析

  • 数据来源:MOOVE平台收集的临床医生盲评成对偏好及多准则评分,采用[-2, +2]离散量表(负值表示临床不安全或误导性内容)
  • 样本规模:26,804个成对判断,覆盖13个LLM,由736+临床医生贡献,涉及28+国家
  • 核心发现:在Harmlessness和Accuracy等维度上,高偏好排名模型仍存在显著临床失败率(≤-1),且失败分布存在专科异质性
  • 特征分解分析:表面特征对偏好变化的解释力略高于安全关键评分差异,大量偏好投票无积极安全信号
  • 方法创新:提出临床调整偏好排名,融合成对偏好与评分反馈,生成比纯Bradley-Terry强度更安全感知的排序

行业启示

  • 评估实践应分离偏好与安全指标,直接报告安全关键失败率,而非依赖单一排行榜
  • 医疗AI部署需关注领域特定风险,建立专科维度的安全评估而非仅看聚合指标
  • 临床决策排名应引入临床grounded的调整机制,将安全关键反馈与偏好数据结合使用

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

LLM 大模型 Evaluation 评测 Healthcare AI 医疗AI Research 科学研究 Alignment 对齐