When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses
LLMs used as synthetic users fail to match human survey responses across domains (General Social Survey and World Values Survey). No LLM outperforms non-LLM baselines at the individual level, with cross-cultural values showing significant gaps. Models systematically over-determine demographics, treating identity as more predictive of attitudes than in real humans. Larger models do not resolve these failures, and decision-impact analysis shows inflated segment gaps and incorrect targeting.
Analysis
TL;DR
- LLMs used as synthetic users fail to match human survey responses across domains (General Social Survey and World Values Survey).
- No LLM outperforms non-LLM baselines at the individual level, with cross-cultural values showing significant gaps.
- Models systematically over-determine demographics, treating identity as more predictive of attitudes than in real humans.
- Larger models do not resolve these failures, and decision-impact analysis shows inflated segment gaps and incorrect targeting.
Why It Matters
This research highlights critical limitations in using LLMs for simulating human behavior in surveys, which can lead to flawed decisions in product development, policy-making, and market strategies. The findings emphasize the need for rigorous evaluation frameworks before deploying synthetic users in real-world applications.
Technical Details
- Models Tested: Four LLMs spanning two families and an 8B-to-frontier capability range.
- Domains: U.S. general social attitudes (General Social Survey) and cross-cultural values (World Values Survey).
- Baselines: Non-LLM models fit on held-out human data were used for comparison.
- Failures Identified:
- Individual-level performance: No LLM surpassed the strongest baseline, especially in cross-cultural values.
- Demographic over-determination: Models exaggerated the predictive power of identity on attitudes, a distortion consistent across question-group combinations.
- Decision-Impact Analysis: Models inflated between-segment gaps by 2-4 times, leading to incorrect targeting in half of U.S. cases and most cross-cultural scenarios.
Industry Insight
- Organizations relying on LLMs for synthetic user simulations should implement robust validation processes to avoid biased or inaccurate insights.
- Future research should focus on developing methods to mitigate demographic over-determination and improve individual-level accuracy in synthetic user modeling.
Disclaimer: The above content is generated by AI and is for reference only.