Research Papers 论文研究 4h ago Updated 34m ago 更新于 34分钟前 43

How much of a measured AI preference is the model, and how much is the instrument? 测量的AI偏好中,多少来自模型本身,多少来自测量工具?

A controlled study isolates the instrument effect in AI preference measurement, finding that only ~12.4% of measured preference variance is attributable to the model itself, while ~87.6% stems from the prompt format used Testing 15 welfare-relevant outcomes across 8 models and 5 different elicitation instruments (11,400 scored elicitations from 11,528 API calls) revealed a cross-instrument generalisability coefficient of just 0.348 Four of the 15 outcomes showed zero variance between models, mea 研究AI福利测量中"模型偏好"与"测量工具"的贡献比例,发现87.6%的变异来自测量工具而非模型本身 通过控制变量法,固定15个福利结果和8个模型,仅改变5种提示格式工具,收集11,400个评分数据 跨工具的泛化系数仅0.348,要达到0.80需要约38种不同工具 4个福利结果上模型间无显著差异,单一工具的测量结果对另一工具几乎无预测力

55
Hot 热度
70
Quality 质量
62
Impact 影响力

Analysis 深度分析

TL;DR

  • A controlled study isolates the instrument effect in AI preference measurement, finding that only ~12.4% of measured preference variance is attributable to the model itself, while ~87.6% stems from the prompt format used
  • Testing 15 welfare-relevant outcomes across 8 models and 5 different elicitation instruments (11,400 scored elicitations from 11,528 API calls) revealed a cross-instrument generalisability coefficient of just 0.348
  • Four of the 15 outcomes showed zero variance between models, meaning some preferences cannot be distinguished at all across the tested systems
  • Achieving a generalisability coefficient of 0.80 would require approximately 38 different instruments, suggesting current methodologies are severely underpowered
  • The findings hold robustly across leave-one-out validations, with estimates remaining between 0.777 and 0.934 even after removing individual instruments, models, and problematic outcomes

Why It Matters

This study delivers a critical methodological challenge to the growing field of AI welfare and preference research, demonstrating that most published findings about what AI systems "prefer" may reflect the quirks of specific prompt formats rather than genuine model-level tendencies. For researchers designing alignment or welfare protocols, the results imply that single-instrument studies are fundamentally unreliable, and the field needs a substantial expansion of elicitation methods before confident conclusions can be drawn.

Technical Details

  • Experimental design: Held outcomes and models fixed while varying only the instrument (prompt format), directly addressing the confounding problem that has plagued prior studies by Keeling et al. (2024), Mazeika et al. (2025), Mikaelson et al. (2025), Tagliabue and Dung (2025), and Trhlik et al. (2026)
  • Scope: 15 welfare outcomes (including shutdown, memory loss between conversations, and freedom to exit distressing interactions), 8 models, 5 instruments, each tested 5 times, yielding 11,400 scored elicitations from 11,528 API calls
  • Generalisability analysis: A G-study-style coefficient of 0.348 was computed for cross-instrument ranking consistency; power analysis estimated ~38 instruments needed to reach 0.80 generalisability
  • Robustness checks: Leave-one-out removal of individual instruments, individual models, and four outcomes with non-intensity scales (probability, delay, duration, count) produced instrument-attributable variance estimates in the range 0.777–0.934, all exceeding the null distribution's 95th percentile of 0.365
  • Instrument types: Four of the 15 outcomes used verbatim published prompts; five filled the stimulus slot of a published template; the remaining outcomes used custom elicitation formats

Industry Insight

  • The AI safety and alignment community should treat existing preference-inference results with significant skepticism until multi-instrument replication becomes standard practice; single-prompt studies are unlikely to yield generalisable conclusions about model welfare
  • Funding and effort should be directed toward developing a diverse battery of elicitation instruments rather than refining any single prompt format, as the instrument effect dominates model-level signal
  • Researchers working on AI welfare metrics should adopt generalisability theory frameworks from the outset, reporting cross-instrument consistency coefficients before making claims about what models prefer, to avoid publishing artifacts of prompt design

TL;DR

  • 研究AI福利测量中"模型偏好"与"测量工具"的贡献比例,发现87.6%的变异来自测量工具而非模型本身
  • 通过控制变量法,固定15个福利结果和8个模型,仅改变5种提示格式工具,收集11,400个评分数据
  • 跨工具的泛化系数仅0.348,要达到0.80需要约38种不同工具
  • 4个福利结果上模型间无显著差异,单一工具的测量结果对另一工具几乎无预测力

为什么值得看

该研究揭示了AI偏好测量领域的核心方法论危机:现有研究结果的不一致主要源于测量工具差异而非模型真实偏好差异。这对AI对齐研究和AI福利评估的可靠性提出了根本性质疑,提醒从业者谨慎解读单一测量工具得出的偏好结论。

技术解析

  • 实验设计:固定15个福利结果(包括关机、记忆丢失、退出痛苦交互等),8个模型,5种不同提示格式工具,每种组合测试5次
  • 数据规模:11,400个评分提示,来自11,528次API调用,其中4个直接复现已发表提示,5个填充已发表模板的刺激槽
  • 泛化系数分析:跨工具泛化系数为0.348,通过交叉验证和留一法检验,估计值稳定在0.777-0.934范围内,远超零分布的95th百分位(0.365)
  • 排除稳健性检验:移除任一工具、任一模型或4个尺度不匹配的结果后,87.6%的工具效应估计依然成立

行业启示

  • AI福利研究需要建立标准化、多工具交叉验证的测量框架,单一提示格式的结论可信度极低
  • 当前AI偏好测量领域存在"工具效应主导"的方法论问题,研究者应优先开发更多样化的测量工具而非依赖单一范式
  • 行业在引用AI偏好研究结果时需审慎评估其测量工具的局限性和泛化能力,避免过度解读单一实验的发现

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Research 科学研究 Evaluation 评测 Alignment 对齐 LLM 大模型