Research Papers 论文研究 7h ago Updated 2h ago 更新于 2小时前 46

VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages VakyArth:评估LLM在印度语言中的语用能力

VakyArth is introduced as the first pragmatic competence benchmark for Indic languages, covering Hindi, Punjabi, Tamil, and Malayalam across five phenomena: deixis, speech acts, implicature, social pragmatics, and coherence. All benchmark items are authored by native speakers and evaluated through multiple-choice questions, natural language inference (NLI), and translation tasks. Multilingual LLMs of varying families and sizes consistently fail on pragmatic meanings rooted in Indic linguistic an VakyArth是首个针对印度语言的语用能力评估基准,覆盖印地语、旁遮普语、泰米尔语和马拉雅拉姆语四种语言 评估涵盖五大语用现象:指示词(deixis)、言语行为(speech acts)、含义(implicature)、社会语用(social pragmatics)和连贯性(coherence) 测试采用多项选择、自然语言推理和翻译三种题型,所有题目均由母语者编写 研究发现现有模型在印度语言文化约定的语用理解上普遍存在系统性缺陷 不同语言类型间存在显著差异:印欧语系语言在翻译任务上优于达罗毗荼语系语言,且自动翻译指标无法准确捕捉语用忠实度

62
Hot 热度
72
Quality 质量
65
Impact 影响力

Analysis 深度分析

TL;DR

  • VakyArth is introduced as the first pragmatic competence benchmark for Indic languages, covering Hindi, Punjabi, Tamil, and Malayalam across five phenomena: deixis, speech acts, implicature, social pragmatics, and coherence.
  • All benchmark items are authored by native speakers and evaluated through multiple-choice questions, natural language inference (NLI), and translation tasks.
  • Multilingual LLMs of varying families and sizes consistently fail on pragmatic meanings rooted in Indic linguistic and cultural conventions.
  • MCQ accuracy consistently exceeds NLI accuracy across all model-language combinations, and translation performance does not reliably track pragmatic understanding.
  • Indo-Aryan languages show a translation advantage over Dravidian languages, and automatic translation metrics miss fluent but pragmatically unfaithful outputs, especially for implicature and deixis.

Why It Matters

This work addresses a critical gap in AI evaluation by moving beyond English-centric pragmatic benchmarks to include linguistically and culturally diverse Indic languages. For AI practitioners building multilingual systems, it highlights that current LLMs lack robust pragmatic competence in low-resource language settings, which has direct implications for deployment in real-world multilingual applications.

Technical Details

  • Benchmark scope: VakyArth covers four Indic languages (Hindi, Punjabi, Tamil, Malayalam) spanning two language families (Indo-Aryan and Dravidian), evaluating five pragmatic phenomena: deixis, speech acts, implicature, social pragmatics, and coherence.
  • Evaluation formats: Three task types are used—multiple-choice questions (MCQ), natural language inference (NLI), and translation—allowing comparative analysis of how different evaluation modalities capture pragmatic understanding.
  • Data provenance: All items are authored by native speakers, ensuring cultural and linguistic authenticity rather than relying on machine-translated or English-derived items.
  • Model evaluation: Multilingual LLMs from varying families and sizes are tested, revealing consistent pragmatic failures across architectures, suggesting the issue is systemic rather than model-specific.
  • Metric limitations: The study demonstrates that automatic translation metrics (e.g., BLEU, COMET) can produce misleadingly high scores for outputs that are fluent yet pragmatically unfaithful, particularly for implicature and deixis.

Industry Insight

  • Multilingual LLM developers should prioritize pragmatic competence evaluation in addition to literal comprehension, as fluency metrics alone are insufficient to guarantee culturally appropriate outputs.
  • The Indo-Aryan vs. Dravidian performance gap suggests that benchmarking and training data curation should account for language-family-specific challenges, not just resource-level differences.
  • Native-speaker-authored benchmarks like VakyArth should become a standard for evaluating non-English LLM capabilities, as machine-translated or English-derived evaluations risk masking pragmatic failures.

TL;DR

  • VakyArth是首个针对印度语言的语用能力评估基准,覆盖印地语、旁遮普语、泰米尔语和马拉雅拉姆语四种语言
  • 评估涵盖五大语用现象:指示词(deixis)、言语行为(speech acts)、含义(implicature)、社会语用(social pragmatics)和连贯性(coherence)
  • 测试采用多项选择、自然语言推理和翻译三种题型,所有题目均由母语者编写
  • 研究发现现有模型在印度语言文化约定的语用理解上普遍存在系统性缺陷
  • 不同语言类型间存在显著差异:印欧语系语言在翻译任务上优于达罗毗荼语系语言,且自动翻译指标无法准确捕捉语用忠实度

为什么值得看

本文填补了印度语言语用能力评估的空白,为多语言大模型的非英语能力诊断提供了首个系统性基准。研究揭示了当前模型在文化约定和隐含意义理解上的系统性缺陷,对开发真正多语言AI具有重要参考价值。

技术解析

  • 基准设计:VakyArth覆盖四种印度语言(印地语、旁遮普语、泰米尔语、马拉雅拉姆语),评估五个语用现象维度,题型包括多项选择、自然语言推理和翻译任务
  • 数据质量:所有测试题目均由母语者编写,确保文化语境和语言习惯的真实性
  • 实验发现:多项选择准确率在所有模型-语言组合中均高于自然语言推理准确率;翻译性能无法可靠反映语用理解水平
  • 语言差异:印欧语系语言(印地语、旁遮普语)在翻译任务上表现出优于达罗毗荼语系语言(泰米尔语、马拉雅拉姆语)的性能
  • 指标局限:自动翻译评估指标会遗漏流畅但语用不忠实的输出,尤其在含义和指示词理解上存在盲区

行业启示

  • 多语言AI开发需重视低资源语言的语用能力评估,不能仅依赖英语基准或字面翻译质量作为能力指标
  • 模型训练应加强对文化约定、隐含意义和社会语境的理解,而非仅关注字面语义
  • 评估体系需要多元化,自动指标(如BLEU、METEOR)无法可靠反映语用忠实度,需结合人工评估和诊断性测试

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

LLM 大模型 Evaluation 评测 Benchmark 基准测试 Dataset 数据集 Research 科学研究