SWORD: Wikidata-based Distortions Reveal Hidden Cross-Lingual Inconsistencies in LLM Factual Error Rejection
SWORD (Systematic Wikidata-based Object-Relation Distortion) is a new benchmark that evaluates LLMs' ability to consistently reject factual errors across eight languages using controlled perturbations of Wikidata triples Models perform better on semantically plausible distortions than nonsensical random substitutions, revealing reliance on distributional familiarity rather than genuine factual verification Cross-lingual performance gaps of up to 28 percentage points (49% relative reduction) emer
Analysis
TL;DR
- SWORD (Systematic Wikidata-based Object-Relation Distortion) is a new benchmark that evaluates LLMs' ability to consistently reject factual errors across eight languages using controlled perturbations of Wikidata triples
- Models perform better on semantically plausible distortions than nonsensical random substitutions, revealing reliance on distributional familiarity rather than genuine factual verification
- Cross-lingual performance gaps of up to 28 percentage points (49% relative reduction) emerge specifically for East Asian languages when models face distorted statements, despite comparable baseline accuracy
- Conventional multilingual benchmarks systematically obscure asymmetric factual reasoning capabilities by aggregating accuracy metrics
Why It Matters
This research exposes a critical blind spot in how multilingual LLM capabilities are evaluated—standard benchmarks reward correct answer selection but fail to test whether models genuinely understand and reject factual errors. For practitioners building multilingual systems, these findings suggest that apparent parity across languages may mask significant vulnerabilities, particularly for East Asian languages where factual error rejection degrades substantially.
Technical Details
- SWORD generates syntactically well-formed but factually incorrect statements in eight widely spoken languages through controlled perturbations of Wikidata triples, including random entity substitutions and semantically plausible property-based selections
- The benchmark specifically measures factual error rejection (the ability to identify and reject false statements) rather than correct answer selection, which is the standard evaluation paradigm
- Evaluation reveals a counterintuitive finding: models achieve higher accuracy on semantically plausible distortions than on nonsensical random substitutions, indicating that surface-level linguistic familiarity can substitute for actual factual verification
- Cross-lingual analysis shows that models with comparable baseline accuracy across languages exhibit dramatic performance drops on East Asian languages under distortion conditions, with gaps reaching 28 percentage points
Industry Insight
- Benchmark design should shift from accuracy-on-correct-answers to error-rejection capabilities, especially for multilingual deployments where asymmetric failures can go undetected
- Organizations deploying LLMs in East Asian markets should conduct rigorous factual consistency audits beyond standard benchmarks, as aggregate metrics may overstate real-world reliability
- The finding that plausible distortions are harder to reject than random ones suggests current training paradigms may need explicit factual verification training rather than relying on distributional patterns alone
Disclaimer: The above content is generated by AI and is for reference only.