TRILOGUE: A Trilingual Spoken Dialogue Fact-Checking Benchmark with Evidence and Paired Audio
TRILOGUE is a large-scale trilingual benchmark for spoken dialogue fact-checking in English, Russian, and Kazakh, addressing the gap where most fact-checking evaluation remains text-only and English-centric The dataset contains nearly 12K dialogues, 187K turns, and 390 hours of paired audio with ASR transcripts and word-level timestamp alignments across all three languages TRILOGUE supports three core tasks: claim check-worthiness detection, source-article evidence retrieval, and claim verificat
Analysis
TL;DR
- TRILOGUE is a large-scale trilingual benchmark for spoken dialogue fact-checking in English, Russian, and Kazakh, addressing the gap where most fact-checking evaluation remains text-only and English-centric
- The dataset contains nearly 12K dialogues, 187K turns, and 390 hours of paired audio with ASR transcripts and word-level timestamp alignments across all three languages
- TRILOGUE supports three core tasks: claim check-worthiness detection, source-article evidence retrieval, and claim verification under claim-only, gold-evidence, and retrieved-evidence conditions
- Baseline results reveal that ASR degradation and cross-lingual transfer remain significant challenges, particularly for Kazakh, while retrieved source evidence substantially narrows the performance gap to gold-evidence verification
Why It Matters
As misinformation increasingly spreads through spoken media—podcasts, interviews, and broadcast dialogue—fact-checking systems must evolve beyond clean written claims to handle the noise and complexity of real-world speech. TRILOGUE provides the first large-scale multilingual benchmark with paired audio, enabling researchers to build and evaluate end-to-end spoken fact-checking pipelines that account for ASR errors, cross-speaker claim distribution, and context-dependent verification.
Technical Details
- Dataset scale and composition: Nearly 12,000 dialogues, 187,000 turns, and 390 hours of paired audio spanning English, Russian, and Kazakh; includes approximately 5,000 human-recorded Russian and Kazakh dialogue files
- Annotations and alignment: Word-level timestamp alignments across all three languages, with turn-level labels for claim verification and source-grounded evidence linking
- Task formulation: Supports claim check-worthiness detection, source-article evidence retrieval, and claim verification under three evidence conditions—claim-only, gold-evidence, and retrieved-evidence—enabling ablation of evidence quality
- Baseline findings: ASR errors significantly degrade verification performance; cross-lingual transfer is especially difficult for Kazakh; retrieved evidence substantially closes the gap relative to gold evidence, highlighting the importance of robust evidence retrieval pipelines
Industry Insight
- The gap between gold-evidence and retrieved-evidence verification suggests that evidence retrieval quality is a critical bottleneck; investing in multilingual retrieval models—particularly for low-resource languages like Kazakh—could yield disproportionate gains
- ASR degradation remains a persistent challenge for spoken fact-checking; systems that jointly model speech and verification, or that are robust to ASR errors, will be better positioned for real-world deployment
- The trilingual design (English, Russian, Kazakh) highlights the need for multilingual benchmarks beyond the typical English-centric evaluation, especially as misinformation ecosystems operate across linguistic boundaries in regions like Central Asia
Disclaimer: The above content is generated by AI and is for reference only.