VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages
VakyArth is introduced as the first pragmatic competence benchmark for Indic languages, covering Hindi, Punjabi, Tamil, and Malayalam across five phenomena: deixis, speech acts, implicature, social pragmatics, and coherence. All benchmark items are authored by native speakers and evaluated through multiple-choice questions, natural language inference (NLI), and translation tasks. Multilingual LLMs of varying families and sizes consistently fail on pragmatic meanings rooted in Indic linguistic an
Analysis
TL;DR
- VakyArth is introduced as the first pragmatic competence benchmark for Indic languages, covering Hindi, Punjabi, Tamil, and Malayalam across five phenomena: deixis, speech acts, implicature, social pragmatics, and coherence.
- All benchmark items are authored by native speakers and evaluated through multiple-choice questions, natural language inference (NLI), and translation tasks.
- Multilingual LLMs of varying families and sizes consistently fail on pragmatic meanings rooted in Indic linguistic and cultural conventions.
- MCQ accuracy consistently exceeds NLI accuracy across all model-language combinations, and translation performance does not reliably track pragmatic understanding.
- Indo-Aryan languages show a translation advantage over Dravidian languages, and automatic translation metrics miss fluent but pragmatically unfaithful outputs, especially for implicature and deixis.
Why It Matters
This work addresses a critical gap in AI evaluation by moving beyond English-centric pragmatic benchmarks to include linguistically and culturally diverse Indic languages. For AI practitioners building multilingual systems, it highlights that current LLMs lack robust pragmatic competence in low-resource language settings, which has direct implications for deployment in real-world multilingual applications.
Technical Details
- Benchmark scope: VakyArth covers four Indic languages (Hindi, Punjabi, Tamil, Malayalam) spanning two language families (Indo-Aryan and Dravidian), evaluating five pragmatic phenomena: deixis, speech acts, implicature, social pragmatics, and coherence.
- Evaluation formats: Three task types are used—multiple-choice questions (MCQ), natural language inference (NLI), and translation—allowing comparative analysis of how different evaluation modalities capture pragmatic understanding.
- Data provenance: All items are authored by native speakers, ensuring cultural and linguistic authenticity rather than relying on machine-translated or English-derived items.
- Model evaluation: Multilingual LLMs from varying families and sizes are tested, revealing consistent pragmatic failures across architectures, suggesting the issue is systemic rather than model-specific.
- Metric limitations: The study demonstrates that automatic translation metrics (e.g., BLEU, COMET) can produce misleadingly high scores for outputs that are fluent yet pragmatically unfaithful, particularly for implicature and deixis.
Industry Insight
- Multilingual LLM developers should prioritize pragmatic competence evaluation in addition to literal comprehension, as fluency metrics alone are insufficient to guarantee culturally appropriate outputs.
- The Indo-Aryan vs. Dravidian performance gap suggests that benchmarking and training data curation should account for language-family-specific challenges, not just resource-level differences.
- Native-speaker-authored benchmarks like VakyArth should become a standard for evaluating non-English LLM capabilities, as machine-translated or English-derived evaluations risk masking pragmatic failures.
Disclaimer: The above content is generated by AI and is for reference only.