Research Papers 论文研究 3h ago Updated 46m ago 更新于 46分钟前 45

Knowing the Form, Not the Function: Automatically Auditing Answer--Authority Decoupling in Legal Benchmarks 知其形不知其用:自动审计法律基准中答案与权威依据的解耦

Legal benchmarks that score only final answers may mask significant authority-grounding failures, as answer correctness and citation accuracy dissociate in both directions Four LLMs spontaneously produced statutory citations across 238 Taiwan bar-examination items without being prompted to do so, enabling automatic joint auditing In criminal law, 24.0–42.4% of responses were answer-correct but missed the gold authority, while 15.2–21.7% were answer-incorrect yet correctly cited it A statutory-re 法律基准测试通常仅评估最终答案正确性,但研究发现答案正确性与法律依据 grounding 存在显著解耦,两者不能互为代理指标 在238个台湾律师考试题目上,4个LLM在未要求引用法规的情况下自发产生权威标记,答案正确与权威引用在两个方向上均出现分离 刑法领域数据显示:24.0-42.4%的有效回答答案正确但遗漏黄金权威,15.2-21.7%答案错误却引用了权威 法规引用具有结构可提取性和外部可验证性,支持自动化联合审计,研究初步扩展至中国大陆民法领域

58
Hot 热度
72
Quality 质量
65
Impact 影响力

Analysis 深度分析

TL;DR

  • Legal benchmarks that score only final answers may mask significant authority-grounding failures, as answer correctness and citation accuracy dissociate in both directions
  • Four LLMs spontaneously produced statutory citations across 238 Taiwan bar-examination items without being prompted to do so, enabling automatic joint auditing
  • In criminal law, 24.0–42.4% of responses were answer-correct but missed the gold authority, while 15.2–21.7% were answer-incorrect yet correctly cited it
  • A statutory-retrieval probe and citation-abstention intervention confirmed that answer and citation behaviors can vary independently at the output level
  • The authors propose joint answer–authority evaluation for statute-grounded legal benchmarks, with preliminary cross-jurisdictional findings from PRC civil law

Why It Matters

This research exposes a critical flaw in how legal AI benchmarks are evaluated: scoring only final answers can falsely inflate model performance by treating authority misses as successes. For AI practitioners building or evaluating legal systems, this means current benchmark scores may not reflect true legal reasoning quality, and joint evaluation frameworks are necessary to ensure models are both correct and properly grounded in statutory authority.

Technical Details

  • The study tested four LLMs on 238 Taiwan bar-examination items under ordinary reasoning prompts that did not request statutory citations, yet models spontaneously produced authority markers
  • Each item has a verified governing provision, enabling automatic joint auditing of answer correctness and authority grounding without manual annotation
  • Two dissociation directions were measured: answer-correct-but-authority-missed (24.0–42.4% in criminal law) and answer-incorrect-but-authority-cited (15.2–21.7%)
  • A statutory-retrieval probe and a permissive citation-abstention intervention were used to demonstrate that answer and citation behaviors can move independently at the output level
  • A preliminary PRC civil-law extension observed similar unrequested authority marking, motivating cross-jurisdictional joint audit frameworks

Industry Insight

  • Benchmark designers should adopt joint answer–authority evaluation metrics rather than relying on answer-only scoring, especially for high-stakes domains like law where grounding is essential
  • The spontaneous citation behavior observed without prompting suggests models have internalized authority-marking tendencies, which evaluators must account for rather than ignore
  • This decoupling phenomenon likely extends beyond legal benchmarks to other domains where factual correctness and source grounding are independently valuable, warranting broader evaluation framework reforms

TL;DR

  • 法律基准测试通常仅评估最终答案正确性,但研究发现答案正确性与法律依据 grounding 存在显著解耦,两者不能互为代理指标
  • 在238个台湾律师考试题目上,4个LLM在未要求引用法规的情况下自发产生权威标记,答案正确与权威引用在两个方向上均出现分离
  • 刑法领域数据显示:24.0-42.4%的有效回答答案正确但遗漏黄金权威,15.2-21.7%答案错误却引用了权威
  • 法规引用具有结构可提取性和外部可验证性,支持自动化联合审计,研究初步扩展至中国大陆民法领域

为什么值得看

本文揭示了当前法律AI基准测试的评估盲区——答案正确性无法代表法律依据的准确性,这对法律垂直领域AI系统的可靠性评估具有重要警示意义。研究提出的联合评估框架为法律基准测试提供了更严谨的评测范式,推动行业从"答案导向"向"依据可追溯"转变。

技术解析

  • 实验设置:使用238个台湾律师考试题目,测试4个LLM在普通推理提示(未要求法规引用)下的表现,每个题目均有经核实的管辖法规作为黄金标准
  • 核心发现:答案正确性与权威 grounding 在两个方向上解耦——刑法中24.0-42.4%的回答答案正确但遗漏黄金权威,15.2-21.7%答案错误却引用了权威
  • 验证实验:通过法规检索探针和允许引用放弃的干预实验,进一步证明答案和引用行为在输出层面可独立变化
  • 扩展研究:初步扩展到中国大陆民法领域,同样观察到未请求引用时的权威标记现象,支持跨法域联合审计的可行性
  • 评估框架:提出联合答案-权威评估方法,利用法规引用的结构可提取性和外部可验证性实现自动化审计

行业启示

  • 基准测试设计:法律AI评测需从单一答案评分转向答案-依据联合评估,避免将答案正确性误认为法律依据可靠性的代理指标
  • 可验证性优先:法规引用作为结构化、可外部验证的信息,应成为法律AI系统评估的核心维度,推动建立可追溯的决策链标准
  • 跨法域标准化:研究在多个法域观察到相似现象,建议推动跨司法管辖区的联合审计框架,建立法律AI基准测试的行业规范

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

LLM 大模型 Evaluation 评测 Benchmark 基准测试 Legal AI 法律AI Research 科学研究