OncoTriad-QA: A Patient-Level Radiology-Pathology-Genomics Benchmark for Pan-Cancer Reasoning
OncoTriad-QA is a novel patient-level benchmark for pan-cancer reasoning that integrates radiology, pathology, genomics, and clinical metadata across 9,281 TCGA cases from 32 cancer cohorts The benchmark contains 86.1k semantic questions with source-grounded LLM-assisted annotations, automated consistency checks, and clinician review OncoVLM, a reference multimodal model, maps radiology, pathology, DNA methylation, and RNA-seq evidence into an LLM interface via learned projectors Existing genera
Analysis
TL;DR
- OncoTriad-QA is a novel patient-level benchmark for pan-cancer reasoning that integrates radiology, pathology, genomics, and clinical metadata across 9,281 TCGA cases from 32 cancer cohorts
- The benchmark contains 86.1k semantic questions with source-grounded LLM-assisted annotations, automated consistency checks, and clinician review
- OncoVLM, a reference multimodal model, maps radiology, pathology, DNA methylation, and RNA-seq evidence into an LLM interface via learned projectors
- Existing general-purpose and medical LLMs show significant limitations on comprehensive pan-cancer QA, especially for cross-modal integration tasks
- Fine-tuned OncoVLM outperforms MedGemma-4B by an average of 10.7 points across MCQ accuracy and BERTScore-F1 under multiple evaluation settings
Why It Matters
This benchmark addresses a critical gap in medical AI evaluation by moving beyond isolated modality tasks to patient-level multi-modal oncology reasoning, which is essential for real-world clinical decision support. The introduction of OncoVLM demonstrates a practical architecture for integrating diverse biomedical data streams into a unified LLM interface, providing a template for future multi-modal medical AI systems.
Technical Details
- Dataset Scale and Composition: 86.1k semantic questions across 9,281 TCGA patient cases spanning 32 cancer cohorts, aligning CT/MRI radiology, whole-slide histopathology, somatic mutations, copy-number alterations, DNA methylation, bulk RNA-seq, and clinical metadata
- Annotation Pipeline: Source-grounded LLM-assisted construction using curated labels, diagnostic reports, molecular profiles, and modality-derived evidence as primary sources of truth, supplemented by automated consistency checks and clinician review
- OncoVLM Architecture: A multimodal model that employs learned projectors to map modality-native radiology, pathology, DNA methylation, and RNA-seq evidence into a unified LLM interface
- Evaluation Metrics: MCQ accuracy and BERTScore-F1, tested under radiology-only, pathology-only, and all-available evidence settings for both multiple-choice and open-ended questions
Industry Insight
- The significant performance gap between existing LLMs and fine-tuned OncoVLM highlights the need for specialized multi-modal training data in oncology, suggesting that general medical LLMs alone are insufficient for comprehensive cancer reasoning tasks
- The source-grounded annotation pipeline with clinician review represents a scalable approach for constructing high-quality multi-modal medical benchmarks that could be adapted to other disease domains
- The learned projector architecture for integrating heterogeneous biomedical data streams offers a practical blueprint for building production-grade multi-modal clinical AI systems that can handle real-world data diversity
Disclaimer: The above content is generated by AI and is for reference only.