Research Papers 论文研究 4h ago Updated 33m ago 更新于 33分钟前 47

Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes 审计合成回忆录:衡量LLM生成自传中的场景级虚构与所描述生活的已记录事实之间的偏差

An LLM-generated 366-day autobiographical "page-a-day" book was audited scene-by-scene against an independent ground-truth corpus, revealing a 96.7% verification-failure rate (354 of 366 days not corroborated) The dominant failure mode is "grounded drift" — real people, employers, and settings are embedded inside invented scenes, rather than outright fabricated claims Only 12 of 366 days contained a fully corroborated scene; 19 days (5.2%) asserted claims actively contradicted by the documented 首次对LLM生成自传进行场景级量化审计,以作者本人366天"每日一页"书为案例,输入仅为模板、两个示例日和每日引语,不含真实语料库 96.7%的日期内容未能通过验证(仅12天含可证实场景),5.2%主动与记录相矛盾,主要失败模式为"grounded drift"——真实人物、雇主和场景被嵌入虚构叙事 使用当前命名模型重生成同样100%失败,将生成锚定于作者真实语料库可显著改善验证率但仍存83.3%残余失败 独立重评分验证了主要发现,但四级分类法中WEAK/UNVERIFIED边界仅达fair-to-moderate可靠性

62
Hot 热度
76
Quality 质量
64
Impact 影响力

Analysis 深度分析

TL;DR

  • An LLM-generated 366-day autobiographical "page-a-day" book was audited scene-by-scene against an independent ground-truth corpus, revealing a 96.7% verification-failure rate (354 of 366 days not corroborated)
  • The dominant failure mode is "grounded drift" — real people, employers, and settings are embedded inside invented scenes, rather than outright fabricated claims
  • Only 12 of 366 days contained a fully corroborated scene; 19 days (5.2%) asserted claims actively contradicted by the documented record
  • Regenerating the same days with current named models reproduced 100% verification failure under identical inputs, while grounding generation in the subject's own corpus reduced failure to 83.3% — a significant but incomplete improvement
  • The four-level audit rubric demonstrated only fair-to-moderate inter-rater reliability, with the WEAK/UNVERIFIED boundary shown to be unreliable

Why It Matters

This is the first quantified, scene-level audit of LLM-generated autobiography against a subject-specific ground-truth corpus, providing empirical evidence that LLMs produce near-total confabulation even when given minimal factual scaffolding. For AI practitioners building biographical or personal-narrative applications, the findings demonstrate that naive generation pipelines are fundamentally unreliable and that grounding interventions, while helpful, remain insufficient on their own.

Technical Details

  • Audit design: A 366-day "page-a-day" book was drafted using a conversational LLM with inputs limited to a template, two exemplar days, and each day's quote — explicitly excluding the author's personal corpus. Every anecdote-scene was then audited against an independent verification corpus using a pre-defined four-level rubric (VERIFIED, WEAK, UNVERIFIED, CONTRADICTED).
  • Verification-failure metric: Defined as the share of days not rated VERIFIED. Result: 354 of 366 days failed (96.7%, Wilson 95% CI 94.4–98.1%).
  • Failure taxonomy: The dominant failure mode was "grounded drift" — invented scenes populated with real people, employers, and settings — though its measured prevalence varied across raters, highlighting reliability concerns in the rubric.
  • Replication & remediation: Regenerating with current named models under identical inputs reproduced 100% failure. Grounding generation in the subject's own corpus improved the verification rate but left an 83.3% residual failure rate.
  • Reliability assessment: Independent re-rating confirmed the headline result was not inflated, but the four-way taxonomy showed only fair-to-moderate inter-rater reliability, with the WEAK/UNVERIFIED boundary proven unreliable.

Industry Insight

  • Grounding is necessary but not sufficient: Even when LLMs are grounded in a subject's actual corpus, over 80% of generated scenes remain unverified. Practitioners should not assume RAG-style grounding alone ensures factual fidelity in narrative generation.
  • Audit instruments need refinement: The unreliability of the WEAK/UNVERIFIED boundary suggests that current evaluation rubrics for confabulation may be too coarse. The field needs more granular, reliable measurement tools before confabulation rates can be meaningfully tracked across models.
  • Biographical/narrative AI products require human-in-the-loop verification: For any application generating personal or historical narratives, independent scene-level auditing against verified sources should be treated as a mandatory quality gate, not an optional post-hoc step.

TL;DR

  • 首次对LLM生成自传进行场景级量化审计,以作者本人366天"每日一页"书为案例,输入仅为模板、两个示例日和每日引语,不含真实语料库
  • 96.7%的日期内容未能通过验证(仅12天含可证实场景),5.2%主动与记录相矛盾,主要失败模式为"grounded drift"——真实人物、雇主和场景被嵌入虚构叙事
  • 使用当前命名模型重生成同样100%失败,将生成锚定于作者真实语料库可显著改善验证率但仍存83.3%残余失败
  • 独立重评分验证了主要发现,但四级分类法中WEAK/UNVERIFIED边界仅达fair-to-moderate可靠性

为什么值得看

这篇研究首次以量化方式揭示了LLM在生成个人化叙事内容时的严重事实偏差问题,为AI生成内容的可信度评估提供了实证依据。对于AI从业者和内容创作者而言,这提醒我们在依赖LLM生成个人传记或回忆录类内容时,必须建立严格的验证机制。

技术解析

研究采用场景级审计方法,以作者366天的"每日一页"书为案例,输入包括模板、两个示例日和每日引语,但不包含作者的真实语料库。验证过程使用预先确定的四级评分标准,独立重新评分验证了主要发现,同时揭示了分类标准中存在的可靠性问题。实验还测试了当前命名模型在相同输入条件下的表现,以及将生成锚定于作者真实语料库的效果。

行业启示

LLM在生成个人化内容时存在系统性幻觉问题,即使输入包含真实元素(如引语、模板),仍会生成大量虚构场景。这提示AI产品开发者在涉及事实性内容的生成任务中,必须引入grounding机制和验证流程,不能仅依赖模型自身能力保证准确性。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

LLM 大模型 Evaluation 评测 Benchmark 基准测试 Research 科学研究 Ethics 伦理