What Reaches Expert Review? Representation, Structural Screening, and Candidate-Form Dependence in AI-Assisted Item Development
Computational evaluators between AI item generation and expert review are not neutral infrastructure—they actively shape which items psychometricians ever evaluate Broad global agreement in semantic geometry concealed consequential local differences: identical wording acquired different construct evidence across embedding configurations Inclusive primary forms across different embedding configurations shared a median of only 6 of 40 items, revealing massive downstream sensitivity to representati
Analysis
TL;DR
- Computational evaluators between AI item generation and expert review are not neutral infrastructure—they actively shape which items psychometricians ever evaluate
- Broad global agreement in semantic geometry concealed consequential local differences: identical wording acquired different construct evidence across embedding configurations
- Inclusive primary forms across different embedding configurations shared a median of only 6 of 40 items, revealing massive downstream sensitivity to representation choices
- Both eligibility policies produced complete forms filling every content cell, yet presented different wording—stability at the aggregate level masked instability at the item level
- The computational evaluator should be treated as an inspectable and revisable component of measurement design rather than a technical preliminary
Why It Matters
This research directly challenges the assumption that AI-assisted psychometric item development pipelines are robust to implementation choices. For AI practitioners building evaluation systems, it demonstrates that seemingly minor decisions about semantic representation and structural screening can dramatically alter which generated content reaches human experts—potentially biasing measurement instruments without any visible warning signs at the aggregate level.
Technical Details
- Two linked in-silico studies tracking 32,000 selected Big Five personality items from semantic representation through structural evaluation to candidate-form construction, using fixed source populations
- Analysis of how different embedding configurations affected construct evidence assignment, item survival rates, and community correspondence metrics across generated source populations
- Comparison of two eligibility policies at the final review boundary, demonstrating that both could fill every content cell in every evaluable form while still producing different item wordings
- Quantification of downstream sensitivity by measuring item overlap across embedding configurations, finding a median overlap of only 6 out of 40 items in inclusive primary forms
- Investigation of the dissociation between global summary stability and local content instability in AI-assisted measurement design pipelines
Industry Insight
- Organizations deploying AI-generated content for high-stakes evaluation should audit not just final outputs but the entire computational pipeline, as representation choices can silently alter measurement properties
- The finding that aggregate completeness masks item-level instability suggests that standard quality metrics (coverage, form completeness) are insufficient—practitioners need granular item-level tracking across pipeline stages
- This work implies that AI-assisted psychometric development should treat computational evaluators as design variables requiring explicit justification and sensitivity analysis, not as fixed technical components
Disclaimer: The above content is generated by AI and is for reference only.