Knowing the Form, Not the Function: Automatically Auditing Answer--Authority Decoupling in Legal Benchmarks
Legal benchmarks that score only final answers may mask significant authority-grounding failures, as answer correctness and citation accuracy dissociate in both directions Four LLMs spontaneously produced statutory citations across 238 Taiwan bar-examination items without being prompted to do so, enabling automatic joint auditing In criminal law, 24.0–42.4% of responses were answer-correct but missed the gold authority, while 15.2–21.7% were answer-incorrect yet correctly cited it A statutory-re
Analysis
TL;DR
- Legal benchmarks that score only final answers may mask significant authority-grounding failures, as answer correctness and citation accuracy dissociate in both directions
- Four LLMs spontaneously produced statutory citations across 238 Taiwan bar-examination items without being prompted to do so, enabling automatic joint auditing
- In criminal law, 24.0–42.4% of responses were answer-correct but missed the gold authority, while 15.2–21.7% were answer-incorrect yet correctly cited it
- A statutory-retrieval probe and citation-abstention intervention confirmed that answer and citation behaviors can vary independently at the output level
- The authors propose joint answer–authority evaluation for statute-grounded legal benchmarks, with preliminary cross-jurisdictional findings from PRC civil law
Why It Matters
This research exposes a critical flaw in how legal AI benchmarks are evaluated: scoring only final answers can falsely inflate model performance by treating authority misses as successes. For AI practitioners building or evaluating legal systems, this means current benchmark scores may not reflect true legal reasoning quality, and joint evaluation frameworks are necessary to ensure models are both correct and properly grounded in statutory authority.
Technical Details
- The study tested four LLMs on 238 Taiwan bar-examination items under ordinary reasoning prompts that did not request statutory citations, yet models spontaneously produced authority markers
- Each item has a verified governing provision, enabling automatic joint auditing of answer correctness and authority grounding without manual annotation
- Two dissociation directions were measured: answer-correct-but-authority-missed (24.0–42.4% in criminal law) and answer-incorrect-but-authority-cited (15.2–21.7%)
- A statutory-retrieval probe and a permissive citation-abstention intervention were used to demonstrate that answer and citation behaviors can move independently at the output level
- A preliminary PRC civil-law extension observed similar unrequested authority marking, motivating cross-jurisdictional joint audit frameworks
Industry Insight
- Benchmark designers should adopt joint answer–authority evaluation metrics rather than relying on answer-only scoring, especially for high-stakes domains like law where grounding is essential
- The spontaneous citation behavior observed without prompting suggests models have internalized authority-marking tendencies, which evaluators must account for rather than ignore
- This decoupling phenomenon likely extends beyond legal benchmarks to other domains where factual correctness and source grounding are independently valuable, warranting broader evaluation framework reforms
Disclaimer: The above content is generated by AI and is for reference only.