How much of a measured AI preference is the model, and how much is the instrument?
A controlled study isolates the instrument effect in AI preference measurement, finding that only ~12.4% of measured preference variance is attributable to the model itself, while ~87.6% stems from the prompt format used Testing 15 welfare-relevant outcomes across 8 models and 5 different elicitation instruments (11,400 scored elicitations from 11,528 API calls) revealed a cross-instrument generalisability coefficient of just 0.348 Four of the 15 outcomes showed zero variance between models, mea
Analysis
TL;DR
- A controlled study isolates the instrument effect in AI preference measurement, finding that only ~12.4% of measured preference variance is attributable to the model itself, while ~87.6% stems from the prompt format used
- Testing 15 welfare-relevant outcomes across 8 models and 5 different elicitation instruments (11,400 scored elicitations from 11,528 API calls) revealed a cross-instrument generalisability coefficient of just 0.348
- Four of the 15 outcomes showed zero variance between models, meaning some preferences cannot be distinguished at all across the tested systems
- Achieving a generalisability coefficient of 0.80 would require approximately 38 different instruments, suggesting current methodologies are severely underpowered
- The findings hold robustly across leave-one-out validations, with estimates remaining between 0.777 and 0.934 even after removing individual instruments, models, and problematic outcomes
Why It Matters
This study delivers a critical methodological challenge to the growing field of AI welfare and preference research, demonstrating that most published findings about what AI systems "prefer" may reflect the quirks of specific prompt formats rather than genuine model-level tendencies. For researchers designing alignment or welfare protocols, the results imply that single-instrument studies are fundamentally unreliable, and the field needs a substantial expansion of elicitation methods before confident conclusions can be drawn.
Technical Details
- Experimental design: Held outcomes and models fixed while varying only the instrument (prompt format), directly addressing the confounding problem that has plagued prior studies by Keeling et al. (2024), Mazeika et al. (2025), Mikaelson et al. (2025), Tagliabue and Dung (2025), and Trhlik et al. (2026)
- Scope: 15 welfare outcomes (including shutdown, memory loss between conversations, and freedom to exit distressing interactions), 8 models, 5 instruments, each tested 5 times, yielding 11,400 scored elicitations from 11,528 API calls
- Generalisability analysis: A G-study-style coefficient of 0.348 was computed for cross-instrument ranking consistency; power analysis estimated ~38 instruments needed to reach 0.80 generalisability
- Robustness checks: Leave-one-out removal of individual instruments, individual models, and four outcomes with non-intensity scales (probability, delay, duration, count) produced instrument-attributable variance estimates in the range 0.777–0.934, all exceeding the null distribution's 95th percentile of 0.365
- Instrument types: Four of the 15 outcomes used verbatim published prompts; five filled the stimulus slot of a published template; the remaining outcomes used custom elicitation formats
Industry Insight
- The AI safety and alignment community should treat existing preference-inference results with significant skepticism until multi-instrument replication becomes standard practice; single-prompt studies are unlikely to yield generalisable conclusions about model welfare
- Funding and effort should be directed toward developing a diverse battery of elicitation instruments rather than refining any single prompt format, as the instrument effect dominates model-level signal
- Researchers working on AI welfare metrics should adopt generalisability theory frameworks from the outset, reporting cross-instrument consistency coefficients before making claims about what models prefer, to avoid publishing artifacts of prompt design
Disclaimer: The above content is generated by AI and is for reference only.