Research Papers 论文研究 5h ago Updated 1h ago 更新于 1小时前 45

TRILOGUE: A Trilingual Spoken Dialogue Fact-Checking Benchmark with Evidence and Paired Audio TRILOGUE:带证据和配对音频的三语口语对话事实核查基准

TRILOGUE is a large-scale trilingual benchmark for spoken dialogue fact-checking in English, Russian, and Kazakh, addressing the gap where most fact-checking evaluation remains text-only and English-centric The dataset contains nearly 12K dialogues, 187K turns, and 390 hours of paired audio with ASR transcripts and word-level timestamp alignments across all three languages TRILOGUE supports three core tasks: claim check-worthiness detection, source-article evidence retrieval, and claim verificat TRILOGUE是首个大规模三语言(英语、俄语、哈萨克语)口语对话事实核查基准,包含近12K对话、187K轮次和390小时配对音频 填补了现有事实核查系统仅评估书面声明、忽视口语对话场景的空白 支持声明检查价值检测、源文章证据检索和声明验证的端到端基准测试 ASR错误退化和跨语言迁移仍是主要挑战,哈萨克语表现尤为突出 检索到的源证据能显著缩小与黄金证据验证的性能差距

58
Hot 热度
72
Quality 质量
65
Impact 影响力

Analysis 深度分析

TL;DR

  • TRILOGUE is a large-scale trilingual benchmark for spoken dialogue fact-checking in English, Russian, and Kazakh, addressing the gap where most fact-checking evaluation remains text-only and English-centric
  • The dataset contains nearly 12K dialogues, 187K turns, and 390 hours of paired audio with ASR transcripts and word-level timestamp alignments across all three languages
  • TRILOGUE supports three core tasks: claim check-worthiness detection, source-article evidence retrieval, and claim verification under claim-only, gold-evidence, and retrieved-evidence conditions
  • Baseline results reveal that ASR degradation and cross-lingual transfer remain significant challenges, particularly for Kazakh, while retrieved source evidence substantially narrows the performance gap to gold-evidence verification

Why It Matters

As misinformation increasingly spreads through spoken media—podcasts, interviews, and broadcast dialogue—fact-checking systems must evolve beyond clean written claims to handle the noise and complexity of real-world speech. TRILOGUE provides the first large-scale multilingual benchmark with paired audio, enabling researchers to build and evaluate end-to-end spoken fact-checking pipelines that account for ASR errors, cross-speaker claim distribution, and context-dependent verification.

Technical Details

  • Dataset scale and composition: Nearly 12,000 dialogues, 187,000 turns, and 390 hours of paired audio spanning English, Russian, and Kazakh; includes approximately 5,000 human-recorded Russian and Kazakh dialogue files
  • Annotations and alignment: Word-level timestamp alignments across all three languages, with turn-level labels for claim verification and source-grounded evidence linking
  • Task formulation: Supports claim check-worthiness detection, source-article evidence retrieval, and claim verification under three evidence conditions—claim-only, gold-evidence, and retrieved-evidence—enabling ablation of evidence quality
  • Baseline findings: ASR errors significantly degrade verification performance; cross-lingual transfer is especially difficult for Kazakh; retrieved evidence substantially closes the gap relative to gold evidence, highlighting the importance of robust evidence retrieval pipelines

Industry Insight

  • The gap between gold-evidence and retrieved-evidence verification suggests that evidence retrieval quality is a critical bottleneck; investing in multilingual retrieval models—particularly for low-resource languages like Kazakh—could yield disproportionate gains
  • ASR degradation remains a persistent challenge for spoken fact-checking; systems that jointly model speech and verification, or that are robust to ASR errors, will be better positioned for real-world deployment
  • The trilingual design (English, Russian, Kazakh) highlights the need for multilingual benchmarks beyond the typical English-centric evaluation, especially as misinformation ecosystems operate across linguistic boundaries in regions like Central Asia

TL;DR

  • TRILOGUE是首个大规模三语言(英语、俄语、哈萨克语)口语对话事实核查基准,包含近12K对话、187K轮次和390小时配对音频
  • 填补了现有事实核查系统仅评估书面声明、忽视口语对话场景的空白
  • 支持声明检查价值检测、源文章证据检索和声明验证的端到端基准测试
  • ASR错误退化和跨语言迁移仍是主要挑战,哈萨克语表现尤为突出
  • 检索到的源证据能显著缩小与黄金证据验证的性能差距

为什么值得看

本文针对现代虚假信息"先被听到而非被阅读"的传播特点,提供了首个多语言口语对话事实核查基准,对推动多模态虚假信息检测研究具有重要价值。对于从事ASR鲁棒性、跨语言NLP和事实核查系统的研究者而言,该数据集为评估端到端系统性能提供了关键资源。

技术解析

  • TRILOGUE数据集包含近12K对话、187K对话轮次和390小时配对音频,覆盖英语、俄语和哈萨克语三种语言,其中近5K俄语和哈萨克语对话为人工录制
  • 数据包含ASR转录文本和词级时间戳对齐,支持三项核心任务:声明检查价值检测、源文章证据检索、声明验证(claim-only/gold-evidence/retrieved-evidence三种输入方式)
  • 基线实验表明ASR退化对事实核查性能有显著负面影响,跨语言迁移在低资源语言(哈萨克语)上表现较差
  • 检索到的源证据能显著缩小与黄金证据验证的性能差距,证明多模态证据检索的有效性

行业启示

  • 随着音频/视频内容成为虚假信息传播的主要载体,事实核查系统必须从纯文本评估扩展到多模态口语对话场景
  • 低资源语言(如哈萨克语)的ASR和NLP基础设施仍需加强,跨语言迁移是提升多语言事实核查能力的关键技术路径
  • 端到端事实核查系统应整合ASR、证据检索和验证模块,而非仅依赖转录文本,以应对真实场景中的ASR错误和上下文依赖问题

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Benchmark 基准测试 Dataset 数据集 Speech 语音 Multimodal 多模态 Evaluation 评测