Research Papers 论文研究 4h ago Updated 30m ago 更新于 30分钟前 43

What Reaches Expert Review? Representation, Structural Screening, and Candidate-Form Dependence in AI-Assisted Item Development 什么能进入专家评审?AI辅助项目开发中的表征、结构筛选与候选形式依赖性

Computational evaluators between AI item generation and expert review are not neutral infrastructure—they actively shape which items psychometricians ever evaluate Broad global agreement in semantic geometry concealed consequential local differences: identical wording acquired different construct evidence across embedding configurations Inclusive primary forms across different embedding configurations shared a median of only 6 of 40 items, revealing massive downstream sensitivity to representati 计算评估器在AI生成项目与专家审查之间起关键筛选作用,其决策直接决定心理测量学家最终接收的内容和证据 语义几何的全局一致性掩盖了重要的局部差异:相同措辞在不同嵌入配置下可获得不同的构念证据 嵌入配置变化导致下游项目选择显著不同,包容性主要形式间仅共享中位数为6/40的项目 全局摘要和完整形式的表观稳定性掩盖了实际到达专家的内容不稳定性 计算评估器应被视为测量设计中可审查和可修订的组成部分,而非中立基础设施

55
Hot 热度
72
Quality 质量
62
Impact 影响力

Analysis 深度分析

TL;DR

  • Computational evaluators between AI item generation and expert review are not neutral infrastructure—they actively shape which items psychometricians ever evaluate
  • Broad global agreement in semantic geometry concealed consequential local differences: identical wording acquired different construct evidence across embedding configurations
  • Inclusive primary forms across different embedding configurations shared a median of only 6 of 40 items, revealing massive downstream sensitivity to representation choices
  • Both eligibility policies produced complete forms filling every content cell, yet presented different wording—stability at the aggregate level masked instability at the item level
  • The computational evaluator should be treated as an inspectable and revisable component of measurement design rather than a technical preliminary

Why It Matters

This research directly challenges the assumption that AI-assisted psychometric item development pipelines are robust to implementation choices. For AI practitioners building evaluation systems, it demonstrates that seemingly minor decisions about semantic representation and structural screening can dramatically alter which generated content reaches human experts—potentially biasing measurement instruments without any visible warning signs at the aggregate level.

Technical Details

  • Two linked in-silico studies tracking 32,000 selected Big Five personality items from semantic representation through structural evaluation to candidate-form construction, using fixed source populations
  • Analysis of how different embedding configurations affected construct evidence assignment, item survival rates, and community correspondence metrics across generated source populations
  • Comparison of two eligibility policies at the final review boundary, demonstrating that both could fill every content cell in every evaluable form while still producing different item wordings
  • Quantification of downstream sensitivity by measuring item overlap across embedding configurations, finding a median overlap of only 6 out of 40 items in inclusive primary forms
  • Investigation of the dissociation between global summary stability and local content instability in AI-assisted measurement design pipelines

Industry Insight

  • Organizations deploying AI-generated content for high-stakes evaluation should audit not just final outputs but the entire computational pipeline, as representation choices can silently alter measurement properties
  • The finding that aggregate completeness masks item-level instability suggests that standard quality metrics (coverage, form completeness) are insufficient—practitioners need granular item-level tracking across pipeline stages
  • This work implies that AI-assisted psychometric development should treat computational evaluators as design variables requiring explicit justification and sensitivity analysis, not as fixed technical components

TL;DR

  • 计算评估器在AI生成项目与专家审查之间起关键筛选作用,其决策直接决定心理测量学家最终接收的内容和证据
  • 语义几何的全局一致性掩盖了重要的局部差异:相同措辞在不同嵌入配置下可获得不同的构念证据
  • 嵌入配置变化导致下游项目选择显著不同,包容性主要形式间仅共享中位数为6/40的项目
  • 全局摘要和完整形式的表观稳定性掩盖了实际到达专家的内容不稳定性
  • 计算评估器应被视为测量设计中可审查和可修订的组成部分,而非中立基础设施

为什么值得看

本文揭示了AI辅助心理测量项目开发流程中一个常被忽视的关键环节:计算评估器的决策如何系统性地影响最终专家审查的内容。对于AI从业者和心理测量学家而言,理解这些筛选机制的敏感性对确保测量工具的有效性和可靠性至关重要。

技术解析

  • 研究设计:两项关联的in-silico研究,追踪32,000个Big Five项目从语义表示到结构评估再到候选形式构建的完整流程,使用固定源群体进行跨配置比较
  • 核心发现:相同措辞在不同嵌入配置下获得不同构念证据,不同项目存活,预期属性可能消失,即使社区对应性改善
  • 量化结果:跨嵌入配置,包容性主要形式仅共享中位数为6/40的项目,反映表示变化通过结构证据和排序产生的总下游后果
  • 政策比较:两种资格政策均填满所有可评估形式的内容单元格,但呈现不同措辞

行业启示

  • AI辅助测量开发中,计算评估器不应被视为技术预处理环节,而应作为测量设计的关键组成部分进行审查和优化
  • 即使全局指标显示稳定,下游内容选择可能对表示方式高度敏感,需建立更细致的局部验证机制
  • 建议建立可追溯的评估流程,记录不同嵌入配置和选择策略对最终项目集的影响,提高AI辅助测量开发的透明度和可重复性

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Evaluation 评测 Research 科学研究 LLM 大模型 Dataset 数据集 Benchmark 基准测试