FIRSTPASS: A Multi-Domain, Multi-Round Peer Review Dataset Grounded in Real Editorial Outcomes
FIRSTPASS is the first large-scale peer review dataset built from complete multi-round editorial dialogues from a multidisciplinary high-impact journal (Nature Communications) It contains 3,668 records spanning five scientific domains: biology, chemistry, neuroscience, physics, and earth science Each record includes the full iterative review cycle: initial referee reports, author point-by-point responses, and updated reviewer assessments Outcome labels (STANDARD for two-round review; EXTENDED fo
Analysis
TL;DR
- FIRSTPASS is the first large-scale peer review dataset built from complete multi-round editorial dialogues from a multidisciplinary high-impact journal (Nature Communications)
- It contains 3,668 records spanning five scientific domains: biology, chemistry, neuroscience, physics, and earth science
- Each record includes the full iterative review cycle: initial referee reports, author point-by-point responses, and updated reviewer assessments
- Outcome labels (STANDARD for two-round review; EXTENDED for three or more rounds) are derived directly from editorial decisions, providing ground truth absent in prior corpora
- Expert reviews average 2,155 words—substantially denser than conference venue reviews—and all data, parsing pipelines, and evaluation scripts are publicly released
Why It Matters
Prior AI peer review models were trained exclusively on Computer Science and Machine Learning conference data, leaving them ill-equipped to evaluate research in other scientific disciplines. FIRSTPASS addresses this critical gap by providing a multidisciplinary, multi-round benchmark grounded in real editorial outcomes, enabling the development and evaluation of AI systems capable of genuine scientific judgment across domains.
Technical Details
- Source: Curated from Nature Communications' mandatory transparent peer review process (instituted November 2022), ensuring editorial authenticity and completeness
- Scale and scope: 3,668 records across five domains (biology, chemistry, neuroscience, physics, earth science), capturing the full iterative structure of scientific validation
- Structure: Each record contains three components—initial referee reports, author point-by-point responses, and updated reviewer assessments—preserving the conversational dynamics of peer review
- Ground truth labels: Editorial outcomes classified as STANDARD (two-round review) or EXTENDED (three or more rounds), providing a supervised signal for predicting review rigor
- Data integrity: An automated audit confirms 100% content integrity; all parsing pipelines, evaluation scripts, and datasets are released for reproducible benchmarking
Industry Insight
- AI systems for scientific peer review must be trained on multidisciplinary data to avoid domain bias; models trained solely on CS/ML venues will fail to recognize discipline-specific validation standards (e.g., contamination controls in biology, NMR spectral assignments in chemistry)
- The multi-round editorial dialogue structure in FIRSTPASS enables research on iterative review simulation, author-response modeling, and dynamic judgment refinement—areas largely unexplored in current AI literature
- The availability of ground-truth outcome labels opens the door to predictive modeling of editorial decisions, which could inform tools for manuscript triage, reviewer assignment, and review quality assessment at publishing houses
Disclaimer: The above content is generated by AI and is for reference only.