Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMs
Current LLM safety alignment training is heavily English-centric, creating a critical cross-lingual safety gap that allows harmful, stereotype-reinforcing outputs in non-English languages The authors introduce INCLUDE, a multilingual evaluation benchmark with 2,604 prompts across six languages (English, Hindi, Bengali, Marathi, Tamil, and Hinglish) to quantify Indian-centric socio-cultural biases Evaluation of ten open- and closed-source LLMs (14,988 bias scores analyzed) reveals Bengali produce
Analysis
TL;DR
- Current LLM safety alignment training is heavily English-centric, creating a critical cross-lingual safety gap that allows harmful, stereotype-reinforcing outputs in non-English languages
- The authors introduce INCLUDE, a multilingual evaluation benchmark with 2,604 prompts across six languages (English, Hindi, Bengali, Marathi, Tamil, and Hinglish) to quantify Indian-centric socio-cultural biases
- Evaluation of ten open- and closed-source LLMs (14,988 bias scores analyzed) reveals Bengali produces the highest average bias in open-source models
- A notable reversal was found for English: it yields the lowest bias in open-source models but the highest bias in closed-source models, suggesting fundamentally different safety alignment approaches
Why It Matters
This research exposes a critical failure mode in deployed AI systems, particularly for voice assistants and spoken dialogue technologies serving linguistically diverse populations like India. As LLMs are increasingly deployed in multilingual contexts, the English-centric safety alignment gap means non-English speakers are disproportionately exposed to harmful biases, making this a pressing concern for AI safety, equity, and responsible deployment at scale.
Technical Details
- INCLUDE Benchmark: A multilingual evaluation framework comprising 2,604 prompts spanning six prompt languages—English, Hindi, Bengali, Marathi, Tamil, and Hinglish (Hindi-English code-mixed)—designed to detect embedded socio-cultural biases specific to Indian contexts
- Model Evaluation: Ten open- and closed-source LLMs were evaluated, generating and analyzing 14,988 individual bias scores across all language-model combinations
- Key Finding on Open-Source Models: Bengali consistently yielded the highest average bias score among open-source models, indicating severe safety alignment gaps for this language
- Key Finding on Closed-Source Models: English demonstrated a reversal pattern, producing the highest bias in closed-source models while yielding the lowest bias in open-source models, suggesting closed-source providers may prioritize non-English safety less aggressively than their open-source counterparts
- Methodology: The study employs statistical analysis of bias scores to quantify the cross-lingual safety gap, framing it as a "safety alignment illusion" where English-centric evaluations create a false sense of comprehensive safety coverage
Industry Insight
- AI developers deploying models in multilingual regions must treat safety alignment as a per-language concern rather than assuming English-centric training generalizes; organizations targeting India and similar markets should adopt benchmarks like INCLUDE to audit their models before deployment
- The English-in-closed-source reversal suggests commercial providers may be optimizing for different safety trade-offs than open-source communities, warranting independent auditing of proprietary models across non-English languages before enterprise adoption
- The rise of code-mixed languages like Hinglish in the benchmark highlights the need for evaluation frameworks that account for real-world linguistic patterns rather than treating each language in isolation, especially in regions with high bilingualism
Disclaimer: The above content is generated by AI and is for reference only.