Research Papers 论文研究 5h ago Updated 46m ago 更新于 46分钟前 43

Reaching the Tail: Calibration Diversity Drives Conformal Coverage under Data Scarcity 触及尾部:校准多样性在数据稀缺下驱动共形覆盖率

Conformal coverage under data scarcity is driven primarily by calibration-set diversity, not rare-event count; support width of the nonconformity-score distribution explains up to 85% of coverage variance versus only 2% for rare-event count A diversity-maximizing calibration set selector improves six-month long-horizon coverage from 67.8% to 81.4%, outperforming Mondrian, shift-robust, and extreme-value alternatives The apparent rare-event threshold previously attributed to Adaptive Conformal In 多时域罕见事件预测在数据稀缺条件下面临核心挑战:标准共形推断的交换性假设被时间序列自相关性违反 控制消融实验揭示:自适应共形推断的"罕见事件阈值"效应实际反映的是校准集大小,非一致性分数分布的支持宽度解释了85%的覆盖率方差,而罕见事件数量仅解释2% 多样性最大化选择器是唯一能改善长期覆盖率的策略,将6个月覆盖率从67.8%提升至81.4%;Mondrian、shift-robust和极值方法均失败,Mondrian在oracle标签下甚至恶化覆盖率 理论命题指出:覆盖率缺口源于校准集上分位数与测试分布的匹配程度,多样性是必要条件但非充分条件 在美国衰退预测框架(RegressorChain两

55
Hot 热度
72
Quality 质量
58
Impact 影响力

Analysis 深度分析

TL;DR

  • Conformal coverage under data scarcity is driven primarily by calibration-set diversity, not rare-event count; support width of the nonconformity-score distribution explains up to 85% of coverage variance versus only 2% for rare-event count
  • A diversity-maximizing calibration set selector improves six-month long-horizon coverage from 67.8% to 81.4%, outperforming Mondrian, shift-robust, and extreme-value alternatives
  • The apparent rare-event threshold previously attributed to Adaptive Conformal Inference is actually an artifact of calibration-set size, as demonstrated through controlled ablation across 200 random calibration sets
  • Coverage deficit is fundamentally determined by how closely the calibration set's upper quantile reaches the test distribution's upper quantile; diversity is necessary but not sufficient
  • The approach is validated on a two-stage U.S. recession-forecasting framework using RegressorChain, with the open question of whether 90% six-month coverage under honest scoring remains unresolved

Why It Matters

This work directly addresses a critical pain point for AI practitioners deploying conformal prediction in real-world settings with limited labeled data, particularly in domains like macroeconomic forecasting where events are inherently rare and autocorrelation violates standard exchangeability assumptions. The finding that calibration diversity—not event rarity—drives coverage has immediate implications for how practitioners select and construct calibration sets, potentially reshaping best practices across uncertainty quantification pipelines.

Technical Details

  • The paper conducts a controlled ablation study across 200 random calibration sets, demonstrating that the support width of the nonconformity-score distribution is the dominant predictor of coverage variance (up to 85%), while rare-event count contributes negligibly (2%)
  • Cross-validation across synthetic conditions and five countries shows consistent results, with Spearman correlations of 0.45–0.66 for support width versus 0.02–0.23 for rare-event count
  • A diversity-maximizing selector is proposed as the calibration set selection strategy, which is the only method tested that meaningfully improves long-horizon conformal coverage
  • Alternative strategies including Mondrian conformal inference, shift-robust methods, and extreme-value theory approaches fail to close the coverage gap; Mondrian notably worsens coverage even under oracle labels
  • The theoretical contribution is a compact proposition formalizing that coverage deficit reflects the proximity of the calibration set's upper quantile to the test distribution's upper quantile
  • Empirical validation uses a two-stage U.S. recession-forecasting framework built on RegressorChain, operating under honest scoring conditions

Industry Insight

  • Practitioners should prioritize calibration set diversity over sheer size or rare-event representation when deploying conformal prediction in data-scarce regimes; this may require rethinking standard calibration set construction protocols
  • The failure of Mondrian and shift-robust methods under these conditions suggests that existing conformal inference extensions may not generalize well to autocorrelated, non-stationary time series with rare events, warranting caution in domains like finance and economics
  • The quantified gap between current 81.4% and the aspirational 90% coverage target provides a concrete benchmark for future research, signaling that while diversity helps, additional mechanisms will be needed to achieve high-confidence long-horizon forecasts under scarcity

TL;DR

  • 多时域罕见事件预测在数据稀缺条件下面临核心挑战:标准共形推断的交换性假设被时间序列自相关性违反
  • 控制消融实验揭示:自适应共形推断的"罕见事件阈值"效应实际反映的是校准集大小,非一致性分数分布的支持宽度解释了85%的覆盖率方差,而罕见事件数量仅解释2%
  • 多样性最大化选择器是唯一能改善长期覆盖率的策略,将6个月覆盖率从67.8%提升至81.4%;Mondrian、shift-robust和极值方法均失败,Mondrian在oracle标签下甚至恶化覆盖率
  • 理论命题指出:覆盖率缺口源于校准集上分位数与测试分布的匹配程度,多样性是必要条件但非充分条件
  • 在美国衰退预测框架(RegressorChain两阶段)中验证,但诚实评分下6个月覆盖率能否达到90%仍是开放问题

为什么值得看

本文揭示了数据稀缺场景下共形推断的核心瓶颈,为宏观经济预测等长序列问题提供了关键理论洞察。研究结果挑战了"罕见事件阈值"的直觉认知,指出校准集多样性比事件数量更重要,对实际部署的不确定性量化系统具有直接指导价值。

技术解析

  • 研究聚焦多时域罕见事件预测,在长宏观经济序列数据约束下,标准不确定性量化方法因自相关性违反交换性假设而失效
  • 跨200个随机校准集的消融实验显示:非一致性分数分布的支持宽度在合成条件和五国数据中均与覆盖率高度相关(Spearman ρ 0.45-0.66),而罕见事件数量相关性极低(ρ 0.02-0.23)
  • 多样性最大化选择器是唯一成功的策略,其他方法(Mondrian分箱、shift-robust、极值理论替代)均未能缩小覆盖率缺口,Mondrian在oracle标签下反而恶化表现
  • 紧凑命题解释机制:覆盖率缺口反映校准集上分位数接近测试分布上分位数的程度,多样性是必要条件但非充分条件
  • 实证验证基于美国衰退预测的两阶段框架(RegressorChain),诚实评分下90%覆盖率目标仍未达成

行业启示

  • 在数据稀缺场景下,应优先优化校准集的样本多样性而非单纯扩充数量,这对金融、医疗等罕见事件预测领域具有直接应用价值
  • 传统共形推断方法在时间序列数据上存在根本性局限, practitioners需重新审视交换性假设的适用性,考虑引入时间感知校正机制
  • 罕见事件预测的研究方向需要调整:现有"阈值效应"解释可能误导资源分配,应聚焦分布匹配质量而非事件计数优化

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Research 科学研究 Evaluation 评测 Training 训练