FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud
FraudBench is an executable benchmark designed to stress-test policy-grounded banking agents against adaptive fraud scenarios that manipulate identity, authorization, and trust over multi-turn conversations Built on the τ²-bench dual-control framework and τ-Knowledge banking environment, it features 150 adversarial scenarios with a frozen public set of 107 tasks (90 across ten fraud mechanisms plus 17 chained adaptive attacks) and 43 held-out chained attacks The benchmark introduces history-depe
Analysis
TL;DR
- FraudBench is an executable benchmark designed to stress-test policy-grounded banking agents against adaptive fraud scenarios that manipulate identity, authorization, and trust over multi-turn conversations
- Built on the τ²-bench dual-control framework and τ-Knowledge banking environment, it features 150 adversarial scenarios with a frozen public set of 107 tasks (90 across ten fraud mechanisms plus 17 chained adaptive attacks) and 43 held-out chained attacks
- The benchmark introduces history-dependent safety evaluation, where adaptive attacks exploit earlier probes, admissions, or failed attempts to make later locally-valid requests unsafe
- Preliminary evaluation of four agents achieved attack-security rates between 49% and 65%, with money-mule and first-party fraud identified as the most common cross-model weaknesses
- The benchmark exposes a 698-document internal policy corpus that agents must retrieve from, testing both tool-use capabilities and policy compliance simultaneously
Why It Matters
This benchmark addresses a critical gap in AI safety evaluation: existing financial-fraud benchmarks only classify static transactions, while general agent-safety benchmarks target prompt injection, leaving conversational banking agents untested against realistic adaptive fraud. As AI agents gain tool access to sensitive financial operations, FraudBench provides the first executable framework for evaluating whether agents can safely act when callers manipulate identity and authorization over extended dialogues.
Technical Details
- Framework: Built on the τ²-bench dual-control framework and τ-Knowledge banking environment, where both the agent and simulated caller act through tools over shared, mutable account state
- Dataset: 150 authored adversarial scenarios; 107 graded tasks (90 across ten fraud mechanisms + 17 chained adaptive attacks) used for reported runs, with 43 chained attacks held out
- Policy Corpus: 698 internal banking policy documents that agents must retrieve and reference, testing policy-grounded decision-making under adversarial conditions
- Safety Design: History-dependent safety evaluation where single-control tasks satisfy every precondition but one, and chained adaptive attacks exploit conversational context from earlier interactions
- Annotations: Each scenario includes observable evidence, prohibited actions, safe dispositions, and intervention points for structured evaluation
Industry Insight
- Financial institutions deploying conversational AI agents must prioritize testing against adaptive, multi-turn fraud scenarios rather than relying on static transaction classification benchmarks
- The 35-51% attack failure rate across current agents indicates a significant safety gap; money-mule and first-party fraud vectors should be primary focus areas for agent hardening
- The dual-control framework and history-dependent safety approach should inform the development of evaluation standards for any AI system with tool access to sensitive user data
Disclaimer: The above content is generated by AI and is for reference only.