Research Papers 论文研究 1d ago Updated 15h ago 更新于 15小时前 45

SWORD: Wikidata-based Distortions Reveal Hidden Cross-Lingual Inconsistencies in LLM Factual Error Rejection SWORD:基于Wikidata的扭曲揭示LLM事实错误拒绝中的隐藏跨语言不一致性

SWORD (Systematic Wikidata-based Object-Relation Distortion) is a new benchmark that evaluates LLMs' ability to consistently reject factual errors across eight languages using controlled perturbations of Wikidata triples Models perform better on semantically plausible distortions than nonsensical random substitutions, revealing reliance on distributional familiarity rather than genuine factual verification Cross-lingual performance gaps of up to 28 percentage points (49% relative reduction) emer 提出SWORD基准,通过Wikidata三元组扰动生成多语言事实错误陈述,评估LLM跨语言事实错误拒绝能力 发现模型在语义合理扰动上准确率高于随机替换,表明依赖分布熟悉度而非真正事实验证 东亚语言在事实错误拒绝任务上性能显著下降,跨语言差距最高达28个百分点(49%相对减少) 揭示多语言事实推理存在不对称能力,聚合准确率指标系统性掩盖了真实缺陷

58
Hot 热度
72
Quality 质量
65
Impact 影响力

Analysis 深度分析

TL;DR

  • SWORD (Systematic Wikidata-based Object-Relation Distortion) is a new benchmark that evaluates LLMs' ability to consistently reject factual errors across eight languages using controlled perturbations of Wikidata triples
  • Models perform better on semantically plausible distortions than nonsensical random substitutions, revealing reliance on distributional familiarity rather than genuine factual verification
  • Cross-lingual performance gaps of up to 28 percentage points (49% relative reduction) emerge specifically for East Asian languages when models face distorted statements, despite comparable baseline accuracy
  • Conventional multilingual benchmarks systematically obscure asymmetric factual reasoning capabilities by aggregating accuracy metrics

Why It Matters

This research exposes a critical blind spot in how multilingual LLM capabilities are evaluated—standard benchmarks reward correct answer selection but fail to test whether models genuinely understand and reject factual errors. For practitioners building multilingual systems, these findings suggest that apparent parity across languages may mask significant vulnerabilities, particularly for East Asian languages where factual error rejection degrades substantially.

Technical Details

  • SWORD generates syntactically well-formed but factually incorrect statements in eight widely spoken languages through controlled perturbations of Wikidata triples, including random entity substitutions and semantically plausible property-based selections
  • The benchmark specifically measures factual error rejection (the ability to identify and reject false statements) rather than correct answer selection, which is the standard evaluation paradigm
  • Evaluation reveals a counterintuitive finding: models achieve higher accuracy on semantically plausible distortions than on nonsensical random substitutions, indicating that surface-level linguistic familiarity can substitute for actual factual verification
  • Cross-lingual analysis shows that models with comparable baseline accuracy across languages exhibit dramatic performance drops on East Asian languages under distortion conditions, with gaps reaching 28 percentage points

Industry Insight

  • Benchmark design should shift from accuracy-on-correct-answers to error-rejection capabilities, especially for multilingual deployments where asymmetric failures can go undetected
  • Organizations deploying LLMs in East Asian markets should conduct rigorous factual consistency audits beyond standard benchmarks, as aggregate metrics may overstate real-world reliability
  • The finding that plausible distortions are harder to reject than random ones suggests current training paradigms may need explicit factual verification training rather than relying on distributional patterns alone

TL;DR

  • 提出SWORD基准,通过Wikidata三元组扰动生成多语言事实错误陈述,评估LLM跨语言事实错误拒绝能力
  • 发现模型在语义合理扰动上准确率高于随机替换,表明依赖分布熟悉度而非真正事实验证
  • 东亚语言在事实错误拒绝任务上性能显著下降,跨语言差距最高达28个百分点(49%相对减少)
  • 揭示多语言事实推理存在不对称能力,聚合准确率指标系统性掩盖了真实缺陷

为什么值得看

现有基准主要奖励模型选择正确答案的能力,而非评估真正的事实理解。SWORD通过扰动式评估揭示了LLM在多语言事实验证中的隐藏缺陷,为构建更可靠的多语言AI系统提供了关键评估工具。

技术解析

  • SWORD(Systematic Wikidata-based Object-Relation Distortion)基准通过控制扰动Wikidata三元组生成语法正确但事实错误的陈述,覆盖8种广泛使用的语言
  • 扰动类型包括随机实体替换和语义合理的属性选择,后者基于语义可塑性生成看似合理但事实错误的语句
  • 实验发现模型在语义合理扰动上的准确率反而高于随机替换,表明模型依赖统计分布熟悉度而非真正的事实验证机制
  • 东亚语言(韩语、日语、中文等)在扰动条件下性能下降最为显著,跨语言性能差距最高达28个百分点

行业启示

  • 多语言LLM评估需超越聚合准确率指标,建立针对事实错误拒绝能力的系统性评估框架
  • 东亚语言模型在事实验证方面存在明显短板,需在训练数据和评估体系上针对性优化
  • 当前多语言AI能力存在不对称性,开发者应警惕"平均性能良好"背后的语言偏见风险

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

LLM 大模型 Benchmark 基准测试 Evaluation 评测 Research 科学研究