Research Papers 论文研究 10h ago Updated 1h ago 更新于 1小时前 53

MyoCardBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios MyoCardBench:用于在临床真实心血管护理场景中评估大型语言模型的现实世界数据基准

MyoCardBench introduces a large-scale, real-world benchmark for evaluating LLMs in cardiovascular care, spanning 13 task-specific datasets with 2,263 items derived from de-identified clinical records. The benchmark assesses models across longitudinal, multimodal, and safety-critical dimensions, revealing significant performance gaps between top models (GPT-5.4) and lower performers (e.g., CardioECGRead at 17.25). Key findings highlight that while GPT-5.4 leads overall, critical weaknesses persis MyoCardBench 是首个覆盖心血管诊疗全流程的真实世界多任务基准,包含2263项来自去标识化临床记录的测试数据。 该基准由16名心内科医师标注并经两位高级专家交叉审核,涵盖诊断、报告生成、伦理决策等13类专科任务。 GPT-5.4在宏观平均分(62.55)和所有维度中排名第一,但整体表现仍远低于人类医生水平;心电图解读与医学伦理任务得分最低(<18)。 开放型任务中“临床完整性”与“关键点覆盖”差距最大,反映模型在复杂情境下的推理与沟通能力存在显著短板。 研究揭示了当前LLM在真实医疗场景中缺乏纵向连续性、多模态整合及安全关键决策能力的核心缺陷。

70
Hot 热度
85
Quality 质量
75
Impact 影响力

Analysis 深度分析

TL;DR

  • MyoCardBench introduces a large-scale, real-world benchmark for evaluating LLMs in cardiovascular care, spanning 13 task-specific datasets with 2,263 items derived from de-identified clinical records.
  • The benchmark assesses models across longitudinal, multimodal, and safety-critical dimensions, revealing significant performance gaps between top models (GPT-5.4) and lower performers (e.g., CardioECGRead at 17.25).
  • Key findings highlight that while GPT-5.4 leads overall, critical weaknesses persist in emergency response, treatment planning, and ethical decision-making tasks, underscoring the need for specialized medical AI development.

Why It Matters

This benchmark is crucial for advancing medical AI by moving beyond isolated exam-style evaluations to simulate real-world clinical workflows. It provides a rigorous framework for identifying model limitations in high-stakes scenarios like emergency rescue and ethics, directly impacting patient safety and trust in AI-assisted care. For researchers, it sets a new standard for developing and validating LLMs in complex, longitudinal healthcare domains.

Technical Details

  • Dataset Composition: MyoCardBench comprises 2,263 items across 13 task-specific datasets, including auxiliary report generation (CardioAuxReport), ECG interpretation (CardioECGRead), communication (CardioComm), emergency rescue (CardioEmergRescue), treatment planning (CardioTreatPlan), and ethical reasoning (CardioEthics).
  • Annotation Process: Data was annotated and reference answers constructed by 16 cardiology physicians, followed by cross-review by two senior cardiologists to ensure clinical accuracy and consistency.
  • Evaluation Protocol: Seven LLMs were tested under standardized zero-shot conditions. Open-ended tasks were assessed using key-point coverage and holistic clinical quality metrics, while CardioEthics was evaluated via accuracy scoring.
  • Performance Metrics: Macro-average scores ranged from 17.25 (CardioECGRead) to 86.38 (CardioAuxReport), with GPT-5.4 achieving the highest macro-average (62.55) and item-weighted mean (62.19). Significant discrepancies were observed between key-point coverage and holistic quality in communication and emergency tasks.

Industry Insight

The results suggest that current LLMs struggle with nuanced, context-dependent clinical tasks such as emergency response and ethical judgment, indicating a need for domain-specific fine-tuning and integration of multimodal data (e.g., imaging, time-series ECG signals). Developers should prioritize improving model robustness in safety-critical areas through targeted training on real-world case studies and continuous validation against expert-reviewed benchmarks like MyoCardBench. Additionally, collaboration between AI engineers and clinicians will be essential to align model outputs with practical healthcare needs and regulatory standards.

TL;DR

  • MyoCardBench 是首个覆盖心血管诊疗全流程的真实世界多任务基准,包含2263项来自去标识化临床记录的测试数据。
  • 该基准由16名心内科医师标注并经两位高级专家交叉审核,涵盖诊断、报告生成、伦理决策等13类专科任务。
  • GPT-5.4在宏观平均分(62.55)和所有维度中排名第一,但整体表现仍远低于人类医生水平;心电图解读与医学伦理任务得分最低(<18)。
  • 开放型任务中“临床完整性”与“关键点覆盖”差距最大,反映模型在复杂情境下的推理与沟通能力存在显著短板。
  • 研究揭示了当前LLM在真实医疗场景中缺乏纵向连续性、多模态整合及安全关键决策能力的核心缺陷。

为什么值得看

本文为AI从业者提供了首个贴近真实临床工作流的心血管领域评估框架,强调从单一知识问答转向端到端诊疗能力评测的必要性。其发现指出主流大模型在高风险医疗场景中的系统性不足,为后续模型优化、安全对齐及医疗AI落地路径提供明确方向。

技术解析

  • 数据集构建:基于13个专项子集共2,263条样本,源自真实脱敏心血管病历与检查报告,覆盖问诊、影像分析、治疗方案制定等全周期环节。
  • 评估体系:采用多维度评分机制——开放题用“关键点覆盖率+整体临床质量”双指标,封闭式如CardioEthics以准确率计,确保评价全面性。
  • 实验设置:7个主流LLM在零-shot条件下生成15,841份输出,避免提示工程干扰,聚焦模型原生能力。
  • 性能分布:CardioAuxReport(辅助报告生成)表现最优(86.38),而CardioECGRead(心电图阅读)与CardioEthics(伦理判断)均低于18分,显示专业细分任务差距巨大。
  • 能力鸿沟:Communication、Emergency Rescue、Treatment Planning三大任务中,“临床完整性”比“知识点覆盖”低超48–52分,说明模型虽能罗列要点却难以组织符合实际诊疗逻辑的方案。

行业启示

  • 医疗AI研发应从“考试导向”转向“流程导向”,构建支持长序列交互、多模态融合及风险管控的综合评估体系。
  • 当前通用大模型在心电图识别、伦理决策等高精度/高责任领域存在严重缺陷,需结合专家规则引擎或强化学习进行垂直加固。
  • 推动建立跨机构协作的真实世界基准库,并引入动态更新机制以匹配临床实践演进,避免评估滞后于技术发展。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

LLM 大模型 Evaluation 评测 Benchmark 基准测试 Healthcare AI 医疗AI