Research Papers 论文研究 5h ago Updated 38m ago 更新于 38分钟前 45

FIRSTPASS: A Multi-Domain, Multi-Round Peer Review Dataset Grounded in Real Editorial Outcomes FIRSTPASS:基于真实编辑决策的多领域多轮同行评审数据集

FIRSTPASS is the first large-scale peer review dataset built from complete multi-round editorial dialogues from a multidisciplinary high-impact journal (Nature Communications) It contains 3,668 records spanning five scientific domains: biology, chemistry, neuroscience, physics, and earth science Each record includes the full iterative review cycle: initial referee reports, author point-by-point responses, and updated reviewer assessments Outcome labels (STANDARD for two-round review; EXTENDED fo 提出FIRSTPASS,首个基于多学科高影响力期刊完整多轮编辑对话的大规模同行评审数据集 数据源自Nature Communications强制透明同行评审(2022年11月起),涵盖生物学、化学、神经科学、物理学和地球科学5个领域 包含3,668条记录,每条记录附带基于编辑决策的结果标签(STANDARD为两轮评审;EXTENDED为三轮及以上),提供此前语料库缺失的真实基准 专家评审平均2,155词,内容密度显著高于会议评审;自动化审计确认100%内容完整性 所有数据、解析管道和评估脚本已开源,支持跨学科AI科学判断的可重复基准测试

58
Hot 热度
72
Quality 质量
65
Impact 影响力

Analysis 深度分析

TL;DR

  • FIRSTPASS is the first large-scale peer review dataset built from complete multi-round editorial dialogues from a multidisciplinary high-impact journal (Nature Communications)
  • It contains 3,668 records spanning five scientific domains: biology, chemistry, neuroscience, physics, and earth science
  • Each record includes the full iterative review cycle: initial referee reports, author point-by-point responses, and updated reviewer assessments
  • Outcome labels (STANDARD for two-round review; EXTENDED for three or more rounds) are derived directly from editorial decisions, providing ground truth absent in prior corpora
  • Expert reviews average 2,155 words—substantially denser than conference venue reviews—and all data, parsing pipelines, and evaluation scripts are publicly released

Why It Matters

Prior AI peer review models were trained exclusively on Computer Science and Machine Learning conference data, leaving them ill-equipped to evaluate research in other scientific disciplines. FIRSTPASS addresses this critical gap by providing a multidisciplinary, multi-round benchmark grounded in real editorial outcomes, enabling the development and evaluation of AI systems capable of genuine scientific judgment across domains.

Technical Details

  • Source: Curated from Nature Communications' mandatory transparent peer review process (instituted November 2022), ensuring editorial authenticity and completeness
  • Scale and scope: 3,668 records across five domains (biology, chemistry, neuroscience, physics, earth science), capturing the full iterative structure of scientific validation
  • Structure: Each record contains three components—initial referee reports, author point-by-point responses, and updated reviewer assessments—preserving the conversational dynamics of peer review
  • Ground truth labels: Editorial outcomes classified as STANDARD (two-round review) or EXTENDED (three or more rounds), providing a supervised signal for predicting review rigor
  • Data integrity: An automated audit confirms 100% content integrity; all parsing pipelines, evaluation scripts, and datasets are released for reproducible benchmarking

Industry Insight

  • AI systems for scientific peer review must be trained on multidisciplinary data to avoid domain bias; models trained solely on CS/ML venues will fail to recognize discipline-specific validation standards (e.g., contamination controls in biology, NMR spectral assignments in chemistry)
  • The multi-round editorial dialogue structure in FIRSTPASS enables research on iterative review simulation, author-response modeling, and dynamic judgment refinement—areas largely unexplored in current AI literature
  • The availability of ground-truth outcome labels opens the door to predictive modeling of editorial decisions, which could inform tools for manuscript triage, reviewer assignment, and review quality assessment at publishing houses

TL;DR

  • 提出FIRSTPASS,首个基于多学科高影响力期刊完整多轮编辑对话的大规模同行评审数据集
  • 数据源自Nature Communications强制透明同行评审(2022年11月起),涵盖生物学、化学、神经科学、物理学和地球科学5个领域
  • 包含3,668条记录,每条记录附带基于编辑决策的结果标签(STANDARD为两轮评审;EXTENDED为三轮及以上),提供此前语料库缺失的真实基准
  • 专家评审平均2,155词,内容密度显著高于会议评审;自动化审计确认100%内容完整性
  • 所有数据、解析管道和评估脚本已开源,支持跨学科AI科学判断的可重复基准测试

为什么值得看

现有AI同行评审数据集局限于计算机科学和机器学习领域,导致模型只能批判消融实验,却从未见过生物学家要求污染控制或化学家质疑NMR谱图分配。FIRSTPASS填补了这一空白,为跨学科AI科学判断的可重复基准测试提供了基础。

技术解析

  • 数据来源:Nature Communications强制透明同行评审(2022年11月实施),确保数据的真实性和完整性
  • 规模与覆盖:3,668条记录,覆盖5个科学领域(生物学、化学、神经科学、物理学、地球科学)
  • 数据结构:完整的多轮迭代结构,包含初始审稿报告、作者逐点回复、更新后的审稿评估
  • 结果标签:基于编辑决策自动标注(STANDARD=两轮评审;EXTENDED=三轮及以上),提供真实ground truth
  • 质量保障:自动化审计确认100%内容完整性,专家评审平均2,155词,显著高于会议评审密度
  • 开源承诺:所有数据、解析管道和评估脚本已发布,支持可重复研究

行业启示

  • 打破领域偏见:当前AI同行评审模型过度集中于CS/ML领域,FIRSTPASS推动跨学科泛化,使AI能理解不同科学领域的评审标准(如生物学的污染控制、化学的NMR谱图分析)
  • 多轮对话价值:完整的多轮迭代结构为研究AI与作者互动、评审改进过程提供了新视角,超越单一静态评审的局限
  • 真实基准建立:编辑决策作为结果标签,为评估AI评审质量提供了可靠基准,推动AI科学判断从"形式批判"向"实质判断"演进

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Dataset 数据集 Benchmark 基准测试 Evaluation 评测 Research 科学研究 LLM 大模型