On the Role of Citations in Preference Data
Humans prefer outputs with more diverse citation sources but fewer total citations, suggesting quality over quantity in attribution LLMs exhibit citation-related preferences even without access to source materials, though these preferences vary significantly across models and datasets The study uses mixed effects models to quantify how citations influence pairwise judgments in scientific question answering Findings have direct implications for designing better preference data collection pipeline
Analysis
TL;DR
- Humans prefer outputs with more diverse citation sources but fewer total citations, suggesting quality over quantity in attribution
- LLMs exhibit citation-related preferences even without access to source materials, though these preferences vary significantly across models and datasets
- The study uses mixed effects models to quantify how citations influence pairwise judgments in scientific question answering
- Findings have direct implications for designing better preference data collection pipelines for reward modeling and post-training
- There is a measurable divergence between human and LLM citation evaluation behavior, highlighting a gap in how machines versus humans assess attribution quality
Why It Matters
This research directly addresses a critical gap in how preference data is constructed for LLM post-training, particularly for tasks requiring attribution and citation. As reward modeling becomes central to aligning LLMs with human values, understanding whether and how models learn to value citations—versus how humans actually do—can prevent misalignment in systems that need to produce trustworthy, verifiable outputs.
Technical Details
- The paper investigates citation preferences among human judges and four open-source LLMs in the context of scientific question answering, using mixed effects statistical models to isolate the influence of citation characteristics on pairwise judgments
- Key citation dimensions analyzed include citation diversity (number of distinct sources) and citation volume (total number of citations), revealing a human preference for breadth over redundancy
- LLMs were evaluated without access to grounding sources, yet still exhibited citation-related biases, suggesting they learn surface-level heuristics rather than semantic understanding of attribution
- The study discusses how these findings should inform preference data collection strategies, particularly for reward models used in RLHF and related post-training methods
Industry Insight
- Teams building reward models should explicitly account for the divergence between human and LLM citation preferences when designing preference datasets, as models may optimize for superficial citation patterns rather than genuine attribution quality
- Citation diversity should be prioritized over raw citation count in evaluation benchmarks and training data, since humans value varied sourcing more than volume
- As attribution becomes increasingly critical for trust and compliance in deployed LLMs, investing in better citation-aware preference data collection now could prevent costly misalignment issues downstream
Disclaimer: The above content is generated by AI and is for reference only.