Research Papers 论文研究 3h ago Updated 1h ago 更新于 1小时前 50

ADAGE: A Language-Agnostic Pipeline for Analogical Reasoning Evaluation ADAGE:一种语言无关的类比推理评估管道

ADAGE is a language-agnostic pipeline for constructing translation-free benchmarks for abstract analogical reasoning, combining native-speaker curation with LLM-assisted generation. The paper validates ADAGE by creating benchmarks for Arabic, Amharic, and Japanese, revealing a significant cultural reasoning gap in multilingual models. Evaluations of 14 open-weight models show that performance on English proverb reasoning does not transfer well to non-English benchmarks, with accuracy drops rangi ADAGE 是一种语言无关的管道,用于构建无需翻译的抽象类比推理基准,结合母语者策划与 LLM 辅助生成。 本文通过为阿拉伯语、阿姆哈拉语和日语创建基准来验证 ADAGE,揭示了多语言模型中存在显著的文化推理差距。 对 14 个开源模型的评估显示,英语谚语推理的性能无法很好地迁移到非英语基准上,准确率下降幅度在 12 至 52 个百分点之间。 作者发布了 ADAGE 管道、三个新基准以及一个评估套件,以鼓励更多基于文化背景的多语言推理研究。

70
Hot 热度
75
Quality 质量
68
Impact 影响力

Analysis 深度分析

TL;DR

  • ADAGE is a language-agnostic pipeline for constructing translation-free benchmarks for abstract analogical reasoning, combining native-speaker curation with LLM-assisted generation.
  • The paper validates ADAGE by creating benchmarks for Arabic, Amharic, and Japanese, revealing a significant cultural reasoning gap in multilingual models.
  • Evaluations of 14 open-weight models show that performance on English proverb reasoning does not transfer well to non-English benchmarks, with accuracy drops ranging from 12 to 52 percentage points.
  • The authors release the ADAGE pipeline, three new benchmarks, and an evaluation suite to encourage more culturally grounded multilingual reasoning research.

Why It Matters

This work addresses a critical flaw in current multilingual AI evaluation: over-reliance on translating English benchmarks, which introduces linguistic artifacts and fails to test true cross-cultural reasoning. By providing a robust framework for generating native-language analogical reasoning tasks, ADAGE enables fairer assessment of model capabilities across diverse languages and cultures, pushing the field toward more equitable and globally applicable AI systems.

Technical Details

  • Pipeline Design: ADAGE integrates human expert curation (native speakers) with large language model assistance to generate high-quality, culturally relevant analogical reasoning examples without translation.
  • Benchmark Construction: Three distinct benchmarks were developed for Arabic, Amharic, and Japanese, focusing on proverbs and idiomatic expressions that reflect local cultural knowledge.
  • Evaluation Setup: 14 open-weight multilingual models were tested on both English and non-English versions of the same analogy tasks to measure performance degradation due to cultural context shifts.
  • Key Finding: Models exhibited substantial declines in accuracy when evaluated on non-English datasets compared to their English counterparts, indicating limited generalization beyond Western-centric training data.
  • Open Release: The full toolkit—including code, dataset samples, and evaluation scripts—is publicly available to facilitate further research into culturally sensitive NLP evaluations.

Industry Insight

AI developers should prioritize building or adopting evaluation frameworks like ADAGE that account for linguistic diversity and cultural nuance rather than relying solely on translated English benchmarks. This shift will help identify blind spots in current models and guide efforts to create more inclusive, globally effective AI products capable of functioning effectively across different regions and communities.

摘要

ADAGE 是一种语言无关的管道,用于构建无需翻译的抽象类比推理基准,结合母语者策划与 LLM 辅助生成。
本文通过为阿拉伯语、阿姆哈拉语和日语创建基准来验证 ADAGE,揭示了多语言模型中存在显著的文化推理差距。
对 14 个开源模型的评估显示,英语谚语推理的性能无法很好地迁移到非英语基准上,准确率下降幅度在 12 至 52 个百分点之间。
作者发布了 ADAGE 管道、三个新基准以及一个评估套件,以鼓励更多基于文化背景的多语言推理研究。

深度分析

TL;DR

  • ADAGE 是一种语言无关的管道,用于构建无需翻译的抽象类比推理基准,结合母语者策划与 LLM 辅助生成。
  • 本文通过为阿拉伯语、阿姆哈拉语和日语创建基准来验证 ADAGE,揭示了多语言模型中存在显著的文化推理差距。
  • 对 14 个开源模型的评估显示,英语谚语推理的性能无法很好地迁移到非英语基准上,准确率下降幅度在 12 至 52 个百分点之间。
  • 作者发布了 ADAGE 管道、三个新基准以及一个评估套件,以鼓励更多基于文化背景的多语言推理研究。

为什么重要

这项工作解决了当前多语言 AI 评估中的一个关键缺陷:过度依赖将英语基准进行翻译,这会引入语言伪影,并无法测试真正的跨文化推理能力。通过提供一种稳健的框架来生成母语类比推理任务,ADAGE 使得在不同语言和文化背景下对模型能力的评估更加公平,推动该领域朝着更公平、更具全球适用性的 AI 系统发展。

技术细节

  • 管道设计:ADAGE 整合了人类专家策划(母语者)与大语言模型协助,在不使用翻译的情况下生成高质量、具有文化相关性的类比推理示例。
  • 基准构建:针对阿拉伯语、阿姆哈拉语和日语分别开发了三个不同的基准,重点关注反映当地文化知识的谚语和习语表达。
  • 评估设置:在相同的类比任务的英文版和非英文版上测试了 14 个开源多语言模型,以测量其性能差异……

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Evaluation 评测 Benchmark 基准测试 Research 科学研究