The Checking Problem: What must be true before AI ships in a regulated firm
Enterprise AI programmes stall because the real bottleneck is not accuracy but the human review burden required before deployment in regulated environments A study of 6 document-heavy workflows across 4 model families and 3 tool configurations (72 total configurations, 5,093 scored outputs) found only 56.1% cleared a strict production bar requiring sustained accuracy, reproducibility, verifiable attribution, and informative confidence signals Tools that provide no confidence score require 100% h
Analysis
TL;DR
- Enterprise AI programmes stall because the real bottleneck is not accuracy but the human review burden required before deployment in regulated environments
- A study of 6 document-heavy workflows across 4 model families and 3 tool configurations (72 total configurations, 5,093 scored outputs) found only 56.1% cleared a strict production bar requiring sustained accuracy, reproducibility, verifiable attribution, and informative confidence signals
- Tools that provide no confidence score require 100% human review of output, while adding source citations and confidence estimates cuts review burden to 49% without sacrificing error tolerance in 17 of 20 configurations
- Self-verification passes reduce review to 44% but cost 2.3x latency and are the only configuration that fails to maintain error tolerance
- The core thesis: AI workflow value is determined less by correctness rate and more by how much a human must still check—a property that is measurable but rarely measured
Why It Matters
This paper directly addresses the gap between AI prototype success and enterprise production failure, particularly in regulated industries like financial services where compliance and auditability are non-negotiable. For AI practitioners, it reframes the evaluation metric from accuracy to review burden, offering a practical lens for deciding which tool configurations are actually deployable. The findings challenge the industry's overemphasis on benchmark accuracy while underinvesting in trustworthiness signals like attribution and confidence calibration.
Technical Details
- Experimental setup: Six document-heavy workflows typical of regulated financial services were evaluated across four model families and three tool configurations, each run three times, yielding 5,093 scored output elements across 72 total configurations
- Dual evaluation bars: Each configuration was assessed against a demonstration bar (single correct run on a single case) and a production bar requiring sustained accuracy, reproducibility across repeats, verifiable attribution, and an informative confidence signal
- Review burden metric: The paper introduces and quantifies "review burden"—the proportion of AI output requiring human inspection—estimated out of sample rather than with hindsight, making it a forward-looking operational metric
- Configuration comparison: No-confidence tools → 100% review; citations + confidence → 49% review (error tolerance held in 17/20 configs); self-verification → 44% review but 2.3x latency cost and failure to hold error tolerance in at least one configuration
- Production survival rate: 57 of 72 configurations cleared the demonstration bar, but only 32 cleared the production bar, revealing a steep drop-off between prototype and deployable performance
Industry Insight
- Organizations should adopt review burden as a first-class deployment metric alongside accuracy, especially in regulated sectors where human-in-the-loop oversight is mandatory; this shifts procurement and evaluation criteria toward configurations that enable triage rather than full manual verification
- The diminishing returns of self-verification suggest that investing in better attribution and calibrated confidence signals is more cost-effective than adding computational overhead, pointing to prompt engineering and output schema design as high-leverage optimization targets
- The 56.1% production survival rate implies that over half of enterprise AI prototypes will fail to ship without explicit investment in reproducibility, verifiability, and confidence calibration—teams should budget for these engineering efforts from the start rather than treating them as post-hoc compliance add-ons
Disclaimer: The above content is generated by AI and is for reference only.