Research Papers 论文研究 3h ago Updated 1h ago 更新于 1小时前 46

When benchmark inferences do not compose: Projectibility in AI evaluation 当基准推理不组合时:AI评估中的可投射性

AI benchmark results rarely translate directly to consequential claims without multiple steps of generalization, interpretation, and extrapolation. Warranted links in evaluation chains do not automatically compose into a warranted chain due to potential mismatches in system, population, outcome, or conditions at interfaces. Projectibility—the extension from observed to unobserved cases—requires careful validation of endpoints, assumptions, dependence, and uncertainty across projections. A non-co 论文指出AI基准测试结果到最终结论的推理链条中,相邻环节的合理支持不能自动保证整体链条的有效性。 提出“不可组合原则”(non-composition principle):只有当端点与假设对齐、依赖性与不确定性被传递时,相邻投影的支持才允许组合。 通过法律研究案例和模拟分析揭示,聚合稳定性可能掩盖后续投影所需的关键区分,导致评估失效。 引入“可投射性审计”(projectibility audit)方法,用于诊断从基准到实际使用论证中的无效连接。 强调在AI评估中需警惕系统、人群、结果或条件在接口处的变化,以及共享数据/模型谱系导致的虚假独立支持。

65
Hot 热度
70
Quality 质量
60
Impact 影响力

Analysis 深度分析

TL;DR

  • AI benchmark results rarely translate directly to consequential claims without multiple steps of generalization, interpretation, and extrapolation.
  • Warranted links in evaluation chains do not automatically compose into a warranted chain due to potential mismatches in system, population, outcome, or conditions at interfaces.
  • Projectibility—the extension from observed to unobserved cases—requires careful validation of endpoints, assumptions, dependence, and uncertainty across projections.
  • A non-composition principle is proposed: adjacent projections are only jointly supported when their endpoints and assumptions align and uncertainty/dependence is properly propagated.
  • Legal-research case studies and simulations demonstrate how independent soundness of individual evaluations (e.g., benchmark vs. deployment) does not guarantee validity of combined arguments.

Why It Matters

This paper addresses a critical epistemic gap in AI evaluation: the assumption that valid intermediate steps in a benchmark-to-deployment argument chain necessarily yield a valid final conclusion. For practitioners and researchers, this means that even if each stage of an evaluation (e.g., model performance on a test set, human review quality, downstream impact assessment) appears individually justified, the overall claim about real-world capability or safety may still be unsupported if the transitions between stages lack proper alignment or uncertainty propagation. The projectibility audit framework offers a practical tool for diagnosing such unsupported joins, which is essential for responsible AI development and regulatory compliance.

Technical Details

  • The paper introduces the concept of "projectibility" as a formal constraint on extending bounded observations (e.g., benchmark scores) to unobserved contexts (e.g., production environments), drawing from Goodman’s problem of rival extensions and argument-based validity frameworks.
  • It proposes a non-composition principle: support for adjacent projections (e.g., from benchmark score to capability claim, then to deployment recommendation) is only valid if (1) endpoints and assumptions align across projections, and (2) dependence structures and uncertainties are explicitly carried through each step.
  • A legal-research case study illustrates how benchmark evidence and a separate deployment study can both be methodologically sound yet remain logically parallel—i.e., neither supports the other’s conclusions without additional bridging assumptions.
  • Reanalysis and simulation experiments show that aggregate stability metrics (e.g., average accuracy across tasks) can obscure task-specific or population-specific distinctions required for later projections, leading to false confidence in generalized claims.
  • The resulting "projectibility audit" provides a structured checklist or workflow to evaluate whether each link in a benchmark-to-use argument chain satisfies the non-composition conditions, identifying unsupported joins where dependencies or mismatches exist.

Industry Insight

AI evaluators and developers should treat benchmark results as starting points rather than endpoints, requiring explicit documentation and validation of each projection step—from raw metrics to capability claims to deployment recommendations. Organizations adopting this framework will need to integrate projectibility audits into their evaluation pipelines, particularly for high-stakes applications like healthcare, finance, or autonomous systems, where misaligned assumptions can lead to catastrophic failures. Additionally, standardization efforts around benchmark reporting should include metadata on target populations, outcome definitions, and dependency structures to enable more rigorous cross-project comparisons and reduce the risk of unwarranted compositional leaps in reasoning.

TL;DR

  • 论文指出AI基准测试结果到最终结论的推理链条中,相邻环节的合理支持不能自动保证整体链条的有效性。
  • 提出“不可组合原则”(non-composition principle):只有当端点与假设对齐、依赖性与不确定性被传递时,相邻投影的支持才允许组合。
  • 通过法律研究案例和模拟分析揭示,聚合稳定性可能掩盖后续投影所需的关键区分,导致评估失效。
  • 引入“可投射性审计”(projectibility audit)方法,用于诊断从基准到实际使用论证中的无效连接。
  • 强调在AI评估中需警惕系统、人群、结果或条件在接口处的变化,以及共享数据/模型谱系导致的虚假独立支持。

为什么值得看

该论文对当前AI评估实践具有深刻批判意义,揭示了主流基准测试向现实部署迁移时的认知陷阱,为构建更严谨的AI能力验证框架提供理论基础。对于从事AI安全、评估标准制定或负责任AI落地的从业者而言,它是理解“性能不等于可用性”这一核心矛盾的关键文献。

技术解析

  • 论文基于Goodman的“新归纳之谜”(rival extensions)构建问题模型,结合argument-based validity架构,形式化检验从观测案例到未观测案例的外推是否 warranted。
  • 核心机制是追踪每个推理步骤中的假设一致性、依赖关系传递及不确定性累积,而非仅关注单步有效性。
  • 使用法律研究作为类比场景,展示即使基准证据与部署研究各自成立,若目标域不匹配则无法形成有效论证链。
  • 通过重分析与模拟实验证明,高聚合稳定性(如平均准确率)可能抹杀关键分布差异,使后续任务外推失去依据。
  • 提出的“可投射性审计”是一种结构化检查流程,用于识别benchmark-to-use路径中缺失的证据断点或未声明的假设跳跃。

行业启示

  • AI评估体系应从单一指标报告转向多阶段可追溯性审计,尤其在将实验室结果推广至真实世界场景前必须执行projectibility check。
  • 开发者与评测机构需主动披露模型训练数据分布、测试环境边界及潜在偏移风险,避免隐含假设误导下游应用决策。
  • 政策制定者在采纳AI系统认证标准时,应强制要求提供跨域泛化能力的实证支撑,而非仅依赖封闭基准得分。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Evaluation 评测 Benchmark 基准测试 Research 科学研究