Research Papers 论文研究 3h ago Updated 1h ago 更新于 1小时前 55

When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses 当合成用户失败:LLM模拟人类调查响应的跨域基准测试

LLMs used as synthetic users fail to match human survey responses across domains (General Social Survey and World Values Survey). No LLM outperforms non-LLM baselines at the individual level, with cross-cultural values showing significant gaps. Models systematically over-determine demographics, treating identity as more predictive of attitudes than in real humans. Larger models do not resolve these failures, and decision-impact analysis shows inflated segment gaps and incorrect targeting. 研究提出评估框架,检验LLM作为“合成用户”模拟人类问卷调查的可靠性。 在跨文化价值观与社会态度两个真实数据集上测试四种模型(8B至前沿能力),发现所有模型均未超越非LLM基线。 两大系统性失败:个体层面预测能力不足、过度依赖人口统计学特征决定态度,且无法通过扩大模型规模解决。 决策影响分析显示,基于LLM模拟的数据会夸大群体间差异2-4倍,导致市场细分错误甚至虚构不存在的群体分裂。 研究呼吁在使用LLM生成调查数据前必须经过严格验证,避免误导产品与政策制定。

72
Hot 热度
85
Quality 质量
78
Impact 影响力

Analysis 深度分析

TL;DR

  • LLMs used as synthetic users fail to match human survey responses across domains (General Social Survey and World Values Survey).
  • No LLM outperforms non-LLM baselines at the individual level, with cross-cultural values showing significant gaps.
  • Models systematically over-determine demographics, treating identity as more predictive of attitudes than in real humans.
  • Larger models do not resolve these failures, and decision-impact analysis shows inflated segment gaps and incorrect targeting.

Why It Matters

This research highlights critical limitations in using LLMs for simulating human behavior in surveys, which can lead to flawed decisions in product development, policy-making, and market strategies. The findings emphasize the need for rigorous evaluation frameworks before deploying synthetic users in real-world applications.

Technical Details

  • Models Tested: Four LLMs spanning two families and an 8B-to-frontier capability range.
  • Domains: U.S. general social attitudes (General Social Survey) and cross-cultural values (World Values Survey).
  • Baselines: Non-LLM models fit on held-out human data were used for comparison.
  • Failures Identified:
    • Individual-level performance: No LLM surpassed the strongest baseline, especially in cross-cultural values.
    • Demographic over-determination: Models exaggerated the predictive power of identity on attitudes, a distortion consistent across question-group combinations.
  • Decision-Impact Analysis: Models inflated between-segment gaps by 2-4 times, leading to incorrect targeting in half of U.S. cases and most cross-cultural scenarios.

Industry Insight

  • Organizations relying on LLMs for synthetic user simulations should implement robust validation processes to avoid biased or inaccurate insights.
  • Future research should focus on developing methods to mitigate demographic over-determination and improve individual-level accuracy in synthetic user modeling.

TL;DR

  • 研究提出评估框架,检验LLM作为“合成用户”模拟人类问卷调查的可靠性。
  • 在跨文化价值观与社会态度两个真实数据集上测试四种模型(8B至前沿能力),发现所有模型均未超越非LLM基线。
  • 两大系统性失败:个体层面预测能力不足、过度依赖人口统计学特征决定态度,且无法通过扩大模型规模解决。
  • 决策影响分析显示,基于LLM模拟的数据会夸大群体间差异2-4倍,导致市场细分错误甚至虚构不存在的群体分裂。
  • 研究呼吁在使用LLM生成调查数据前必须经过严格验证,避免误导产品与政策制定。

为什么值得看

该研究揭示了当前LLM在社会科学模拟中的根本性缺陷,对依赖AI进行市场调研、公共政策建模或用户行为预测的团队具有警示意义——盲目信任合成数据可能导致严重误判。其提供的跨域基准与评估方法可被直接用于内部审计自身系统的风险边界。

技术解析

  • 实验设计采用统一协议,在General Social Survey(美国社会态度)和World Values Survey(跨文化价值观)两个独立真实数据集上运行,确保结果泛化性。
  • 对比对象包括四种不同家族与规模的LLM(从8B参数到前沿模型),以及基于保留人类数据训练的非LLM统计基线(如逻辑回归等)。
  • 核心发现之一:所有LLM在个体响应预测上均弱于最强基线,尤其在跨文化维度差距显著,且该差距经距离感知评分与正确性校准后仍成立。
  • 另一关键问题:模型将人口变量(如年龄、性别、教育)视为比实际更强有力的态度预测因子,这种“人口决定论偏差”在几乎所有题目组中一致存在,并通过编码不变量度量验证其鲁棒性。
  • 增大模型容量或提升能力等级未能缓解上述两类失败,表明问题源于架构或训练机制而非单纯规模限制。

行业启示

  • 企业在部署LLM驱动的用户调研系统时,应引入类似本文提出的“双失败检测机制”,优先排查是否存在人口过度拟合与个体预测失效现象。
  • 政策制定者若使用AI模拟公众意见,需警惕由此产生的虚假群体分化风险,建议结合真实抽样数据进行交叉校验后再作决策。
  • 未来方向应转向构建混合式合成用户系统——融合规则引擎、因果推断与小样本人类反馈,而非单纯依赖大语言模型的自回归生成能力。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

LLM 大模型 Evaluation 评测 Benchmark 基准测试 Research 科学研究