Research Papers 论文研究 7h ago Updated 2h ago 更新于 2小时前 46

Leveraging Large Language Models for Systematic Literature Review of Disease Spread Models 利用大语言模型进行疾病传播模型的系统性文献综述

An LLM pipeline was developed to extract model-relevant information from 536 peer-reviewed agent-based modeling papers on disease spread GPT-4.1 achieved ~77.95% paper-level accuracy and GPT-5.0 achieved ~81.67% paper-level accuracy when compared against human-conducted SLR Field-level accuracy varied widely (32.40% to 100.00%), with complex or subjective fields performing less reliably LLM inter-agreement was identified as a quality indicator: low agreement signals hallucinations, while high ag 研究开发了LLM流水线,从536篇基于智能体的建模论文中提取疾病传播模型相关信息,并与人工系统综述结果对比验证 GPT-4.1论文级准确率达77.95%,GPT-5.0提升至81.67%,但字段级准确率波动显著(32.40%-100%) 发现LLM间一致性可作为输出质量指标:低一致性提示幻觉风险,高一致性结合低准确率则可能反映人工数据集存在噪声或错误 复杂或主观性字段提取可靠性较低,提示词工程对性能提升具有关键作用 研究系统阐述了LLM在建模与仿真领域系统文献综述中的应用潜力与现实局限

62
Hot 热度
72
Quality 质量
65
Impact 影响力

Analysis 深度分析

TL;DR

  • An LLM pipeline was developed to extract model-relevant information from 536 peer-reviewed agent-based modeling papers on disease spread
  • GPT-4.1 achieved ~77.95% paper-level accuracy and GPT-5.0 achieved ~81.67% paper-level accuracy when compared against human-conducted SLR
  • Field-level accuracy varied widely (32.40% to 100.00%), with complex or subjective fields performing less reliably
  • LLM inter-agreement was identified as a quality indicator: low agreement signals hallucinations, while high agreement with low accuracy may indicate noise in the human reference dataset
  • The study provides practical insights into prompt development and outlines both the potential and limitations of automating full-scale systematic literature reviews in modeling and simulation domains

Why It Matters

This work directly addresses a critical bottleneck in AI and computational research: the time-intensive nature of systematic literature reviews. As LLMs become more capable, demonstrating their reliability (and limitations) for structured information extraction from scientific literature helps practitioners decide when to trust automated pipelines versus human curation. The finding that LLM agreement can serve as a proxy for output quality offers a practical, deployable heuristic for researchers building review automation systems.

Technical Details

  • Dataset: 536 peer-reviewed agent-based modeling papers focused on disease spread, with ground truth established by a human-conducted systematic literature review
  • Models evaluated: GPT-4.1 and GPT-5.0, tested on an LLM pipeline designed for structured information extraction from scientific papers
  • Metrics: Paper-level accuracy (77.95% for GPT-4.1, 81.67% for GPT-5.0) and field-level accuracy (range: 32.40%–100.00%), with performance varying by field complexity and subjectivity
  • Key methodological insight: Cross-model agreement was used as a diagnostic signal — low LLM-LLM agreement correlates with hallucinations, while high agreement paired with low accuracy suggests potential errors or noise in the human-annotated reference dataset rather than model failure
  • Domain: Applied to agent-based modeling in epidemiology, within the broader context of systematic literature review automation

Industry Insight

  • LLMs are approaching human-level reliability for structured literature extraction on well-defined, objective fields, but remain significantly less trustworthy on subjective or complex classification tasks — practitioners should implement human-in-the-loop review for high-stakes fields
  • The cross-model agreement diagnostic is a deployable quality-control mechanism: teams building automated review pipelines can use inter-model disagreement as an early warning signal for hallucination-prone outputs without requiring full human re-review
  • Prompt engineering and field-level calibration remain essential; the wide accuracy range (32–100%) indicates that one-size-fits-all LLM pipelines will underperform, and domain-specific prompt optimization is necessary before scaling to full literature review automation

TL;DR

  • 研究开发了LLM流水线,从536篇基于智能体的建模论文中提取疾病传播模型相关信息,并与人工系统综述结果对比验证
  • GPT-4.1论文级准确率达77.95%,GPT-5.0提升至81.67%,但字段级准确率波动显著(32.40%-100%)
  • 发现LLM间一致性可作为输出质量指标:低一致性提示幻觉风险,高一致性结合低准确率则可能反映人工数据集存在噪声或错误
  • 复杂或主观性字段提取可靠性较低,提示词工程对性能提升具有关键作用
  • 研究系统阐述了LLM在建模与仿真领域系统文献综述中的应用潜力与现实局限

为什么值得看

本研究为AI辅助学术研究提供了实证依据,展示了大语言模型在自动化文献综述中的实际应用价值与能力边界。对AI从业者和研究人员而言,它揭示了多模型一致性验证作为质量控制手段的可行性,为构建可靠的AI科研工具提供了重要参考。

技术解析

  • 研究构建了专门的LLM信息提取流水线,处理536篇同行评审的基于智能体的建模论文,提取与疾病传播模型相关的关键信息,并与人工 conducted 的系统综述进行对比评估
  • GPT-4.1和GPT-5.0在论文级分类任务上分别达到77.95%和81.67%的准确率,但字段级准确率差异显著(32.40%-100%),复杂或主观性字段的表现明显较差
  • 提出多模型一致性作为质量评估指标的新方法:低一致性通常意味着模型幻觉,而高一致性配合低准确率则可能指向人工标注数据本身存在噪声或错误
  • 研究为提示词开发提供了实践性见解,并系统性地讨论了LLM在建模和仿真领域进行全面系统综述的潜力与局限性

行业启示

  • LLM辅助文献综述已从概念验证走向实际应用阶段,但需要建立多模型交叉验证机制来识别和降低幻觉风险,确保研究结果的可靠性
  • 专业领域的信息提取任务对模型能力提出更高要求,复杂字段和主观判断仍是当前技术的薄弱环节,需针对性优化提示策略
  • 研究数据质量评估需要新的方法论,多模型一致性分析为区分模型幻觉与人工标注噪声提供了可行路径,对AI科研工具开发具有指导意义

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

LLM 大模型 Research 科学研究 Healthcare AI 医疗AI