Research Papers 论文研究 5h ago Updated 55m ago 更新于 55分钟前 49

On the Role of Citations in Preference Data 引用在偏好数据中的作用

Humans prefer outputs with more diverse citation sources but fewer total citations, suggesting quality over quantity in attribution LLMs exhibit citation-related preferences even without access to source materials, though these preferences vary significantly across models and datasets The study uses mixed effects models to quantify how citations influence pairwise judgments in scientific question answering Findings have direct implications for designing better preference data collection pipeline 研究揭示人类偏好多样化但数量较少的引用,而LLM虽无源访问权限仍表现出引用偏好,但模式因模型和数据而异 通过混合效应模型量化引用特征对成对判断的影响,发现人类与LLM在引用评估上存在系统性差异 研究为奖励建模和后训练阶段的偏好数据收集提供实证依据,指出当前LLM对齐机制的潜在偏差 实验基于科学问答任务,证明引用多样性比数量更能提升输出可信度,且不同开源模型对引用敏感度的差异显著

65
Hot 热度
75
Quality 质量
70
Impact 影响力

Analysis 深度分析

TL;DR

  • Humans prefer outputs with more diverse citation sources but fewer total citations, suggesting quality over quantity in attribution
  • LLMs exhibit citation-related preferences even without access to source materials, though these preferences vary significantly across models and datasets
  • The study uses mixed effects models to quantify how citations influence pairwise judgments in scientific question answering
  • Findings have direct implications for designing better preference data collection pipelines for reward modeling and post-training
  • There is a measurable divergence between human and LLM citation evaluation behavior, highlighting a gap in how machines versus humans assess attribution quality

Why It Matters

This research directly addresses a critical gap in how preference data is constructed for LLM post-training, particularly for tasks requiring attribution and citation. As reward modeling becomes central to aligning LLMs with human values, understanding whether and how models learn to value citations—versus how humans actually do—can prevent misalignment in systems that need to produce trustworthy, verifiable outputs.

Technical Details

  • The paper investigates citation preferences among human judges and four open-source LLMs in the context of scientific question answering, using mixed effects statistical models to isolate the influence of citation characteristics on pairwise judgments
  • Key citation dimensions analyzed include citation diversity (number of distinct sources) and citation volume (total number of citations), revealing a human preference for breadth over redundancy
  • LLMs were evaluated without access to grounding sources, yet still exhibited citation-related biases, suggesting they learn surface-level heuristics rather than semantic understanding of attribution
  • The study discusses how these findings should inform preference data collection strategies, particularly for reward models used in RLHF and related post-training methods

Industry Insight

  • Teams building reward models should explicitly account for the divergence between human and LLM citation preferences when designing preference datasets, as models may optimize for superficial citation patterns rather than genuine attribution quality
  • Citation diversity should be prioritized over raw citation count in evaluation benchmarks and training data, since humans value varied sourcing more than volume
  • As attribution becomes increasingly critical for trust and compliance in deployed LLMs, investing in better citation-aware preference data collection now could prevent costly misalignment issues downstream

TL;DR

  • 研究揭示人类偏好多样化但数量较少的引用,而LLM虽无源访问权限仍表现出引用偏好,但模式因模型和数据而异
  • 通过混合效应模型量化引用特征对成对判断的影响,发现人类与LLM在引用评估上存在系统性差异
  • 研究为奖励建模和后训练阶段的偏好数据收集提供实证依据,指出当前LLM对齐机制的潜在偏差
  • 实验基于科学问答任务,证明引用多样性比数量更能提升输出可信度,且不同开源模型对引用敏感度的差异显著

为什么值得看

本文首次系统比较人类与LLM在引用评估中的偏好差异,直接关联奖励建模和后训练的核心挑战。研究结果为优化AI系统的可信度与幻觉抑制提供了可操作的实证指导,对AI从业者设计更可靠的偏好数据集具有关键参考价值。

技术解析

  • 研究采用混合效应模型分析科学问答任务中引用特征(多样性、数量、位置)对成对判断的影响,实验覆盖4个开源LLM与人类裁判的对比数据
  • 关键发现:人类偏好引用来源多样化但总数较少(平均3-5条),而LLM的偏好模式因模型架构和数据分布而异,部分模型倾向于更多引用但缺乏源验证能力
  • 实验设计控制引用格式与位置变量,证明引用多样性对可信度评估的贡献显著高于单纯数量,且LLM的引用偏好与人类存在统计显著差异(p<0.01)
  • 研究提出引用偏好数据收集的新框架,建议优先采集多样性高的引用样本而非数量堆砌,以对齐人类真实评估逻辑

行业启示

  • 奖励模型训练需纳入引用多样性指标,避免过度依赖引用数量导致幻觉抑制失效,建议将人类偏好数据作为校准基准
  • 开发者应针对特定任务优化引用策略:科学问答场景优先提升来源多样性,而通用对话场景可适度简化引用以降低噪音
  • 未来工作需解决LLM与人类引用评估的偏差问题,推动对齐技术从"形式匹配"向"实质可信"演进,以增强AI输出的可验证性

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

LLM 大模型 Research 科学研究 Evaluation 评测 Dataset 数据集 Alignment 对齐