Research Papers 论文研究 2d ago Updated 1d ago 更新于 1天前 49

LongNovel: A Multi-Scale Benchmark for Hallucination Detection in Long-Context Novel Summarization LongNovel:用于长上下文小说摘要幻觉检测的多尺度基准

LongNovel is a multi-scale bilingual (Chinese-English) benchmark for detecting hallucinations in long-context novel summarization, addressing a gap in existing research The benchmark uses 29 Chinese novels ranging from 16k to 100k tokens plus chapter-level data from BookSum, enabling analysis of how hallucinations scale with context length Eight hallucination types were designed, with data quality ensured through Multi-Model Arbitration, Entity-Referenced Hallucination Generation, and manual tes 提出LongNovel,首个面向长上下文小说摘要的多尺度双语幻觉检测基准 基于29部中文小说(16k-100k token)和BookSum数据集构建,覆盖多长度尺度 设计8种幻觉类型,结合多模型仲裁与实体引用幻觉生成技术确保数据质量 实验验证LongNovel具有挑战性,为长上下文幻觉研究提供可靠评估平台

65
Hot 热度
75
Quality 质量
70
Impact 影响力

Analysis 深度分析

TL;DR

  • LongNovel is a multi-scale bilingual (Chinese-English) benchmark for detecting hallucinations in long-context novel summarization, addressing a gap in existing research
  • The benchmark uses 29 Chinese novels ranging from 16k to 100k tokens plus chapter-level data from BookSum, enabling analysis of how hallucinations scale with context length
  • Eight hallucination types were designed, with data quality ensured through Multi-Model Arbitration, Entity-Referenced Hallucination Generation, and manual test-set revision
  • Experimental results confirm LongNovel is a challenging benchmark, and the dataset has been released for community use

Why It Matters

As context windows continue to expand, hallucination detection in long-context summarization remains a critical bottleneck for deploying LLMs in real-world literary and narrative tasks. LongNovel provides the first multi-scale benchmark that systematically evaluates how hallucination rates evolve as context grows, offering researchers a structured way to measure and improve long-context reliability.

Technical Details

  • Dataset construction: 29 Chinese novels (16k–100k tokens each) combined with chapter-level summaries from the BookSum dataset, creating a bilingual benchmark covering diverse narrative complexities
  • Hallucination taxonomy: 8 distinct hallucination types were designed to categorize different failure modes in long-context summarization, enabling fine-grained analysis
  • Data generation pipeline: Combines Multi-Model Arbitration (using multiple models to cross-validate outputs) and Entity-Referenced Hallucination Generation (injecting hallucinations tied to specific entities) to ensure both authenticity and balanced category distribution
  • Quality assurance: Manual revision of the test set to guarantee data reliability and eliminate artifacts from automated generation
  • Multi-scale design: The varying novel lengths (16k to 100k tokens) allow researchers to study hallucination patterns across different context lengths rather than at a single fixed scale

Industry Insight

  • Benchmark designers should prioritize multi-scale evaluation: single-length benchmarks mask critical failure modes that only emerge at extreme context lengths, so future benchmarks should span a wide range of token counts
  • The combination of entity-referenced hallucination injection with multi-model arbitration offers a replicable template for constructing high-quality hallucination datasets in other domains beyond novel summarization
  • As long-context models become commodity, hallucination detection will shift from a research curiosity to a production-critical requirement; investing in robust evaluation benchmarks now will differentiate teams that ship reliable long-context applications

TL;DR

  • 提出LongNovel,首个面向长上下文小说摘要的多尺度双语幻觉检测基准
  • 基于29部中文小说(16k-100k token)和BookSum数据集构建,覆盖多长度尺度
  • 设计8种幻觉类型,结合多模型仲裁与实体引用幻觉生成技术确保数据质量
  • 实验验证LongNovel具有挑战性,为长上下文幻觉研究提供可靠评估平台

为什么值得看

随着大模型上下文窗口不断扩展,长文本摘要中的幻觉问题日益突出。LongNovel填补了长上下文小说摘要幻觉检测基准的空白,为评估和改进模型在复杂叙事场景下的事实一致性提供了重要工具。

技术解析

  • 数据集构建:从29部中文小说(16k-100k token)和BookSum数据集提取章节级数据,形成多尺度测试集,覆盖从短篇到长篇的不同上下文长度。
  • 幻觉类型设计:定义8种幻觉类型,通过多模型仲裁和实体引用幻觉生成技术,确保幻觉类别的平衡分布和数据真实性。
  • 数据质量控制:测试集内容经过人工修订,保证数据可靠性,同时保持幻觉分布的多样性。
  • 基准评估:实验结果表明LongNovel对当前模型具有挑战性,为幻觉检测研究提供标准化评估平台。

行业启示

  • 长上下文能力已成为大模型竞争焦点,幻觉检测基准的建立将推动模型在事实一致性方面的改进。
  • 小说摘要场景比新闻或论文更复杂,需要更细致的幻觉分类和评估方法,这为垂直领域应用提供参考。
  • 多尺度基准设计思路可推广至其他长文本任务,如代码生成、法律文档分析等场景的评估体系建设。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Benchmark 基准测试 Evaluation 评测 Dataset 数据集 Research 科学研究