Research Papers 论文研究 5h ago Updated 1h ago 更新于 1小时前 49

Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culture Terminal-Bench-LILT:基于语言、地区和文化的多语言智能体编码基准

Terminal-Bench-LILT introduces 300 authentic multilingual coding tasks across ten languages (Arabic, Czech, German, Spanish, Hindi, Japanese, Korean, Serbian, Turkish, and Chinese), addressing a critical gap in non-English coding evaluation. Tasks target region-specific and culture-grounded software development issues such as internationalization, encoding, text normalization, and cultural conventions that lack direct English equivalents. All tasks are authored by native-speaker programmers and 提出Terminal-Bench-LILT基准,包含300个多语言编码任务,覆盖阿拉伯语、捷克语、德语、西班牙语、印地语、日语、韩语、塞尔维亚语、土耳其语和中文 任务聚焦非英语软件开发特有难题,如国际化、编码处理、文本归一化和文化惯例,无直接英文对应 六个前沿模型评估显示最强模型仅达63.1%通过率,大量任务未被任何模型解决 多语言编码能力与通用编程基准排名无关,是独立且未被充分探索的能力维度

65
Hot 热度
75
Quality 质量
70
Impact 影响力

Analysis 深度分析

TL;DR

  • Terminal-Bench-LILT introduces 300 authentic multilingual coding tasks across ten languages (Arabic, Czech, German, Spanish, Hindi, Japanese, Korean, Serbian, Turkish, and Chinese), addressing a critical gap in non-English coding evaluation.
  • Tasks target region-specific and culture-grounded software development issues such as internationalization, encoding, text normalization, and cultural conventions that lack direct English equivalents.
  • All tasks are authored by native-speaker programmers and validated through a multi-stage quality control pipeline, ensuring authenticity and linguistic accuracy.
  • Evaluation of six frontier models shows the strongest model achieves only 63.1% pass rate, with many tasks unsolved by any model, revealing significant multilingual coding gaps.
  • Performance varies substantially across languages and does not correlate with general coding benchmark rankings, establishing multilingual coding competence as a distinct capability axis.

Why It Matters

This benchmark exposes a critical blind spot in AI coding agent evaluation: the overwhelming focus on English-language tasks masks severe deficiencies in multilingual software development capabilities. For practitioners deploying coding agents globally, these findings signal that high performance on standard English benchmarks does not translate to real-world multilingual environments, necessitating new evaluation frameworks and training strategies.

Technical Details

  • Benchmark composition: 300 coding tasks distributed across ten languages, each targeting non-English-specific development challenges including internationalization (i18n), character encoding, text normalization, locale-aware formatting, and cultural conventions.
  • Authorship and validation: Tasks are authored by native-speaker programmers and undergo a multi-stage quality control pipeline to ensure linguistic authenticity and technical correctness.
  • Evaluation scope: Six frontier AI coding models were tested, with pass rates measured per language and overall, revealing substantial performance variation across languages.
  • Key finding: The best-performing model achieved only 63.1% pass rate, and numerous tasks were unsolved by all models, indicating that current frontier agents lack robust multilingual coding competence.
  • Decoupling from general coding ability: Performance on Terminal-Bench-LILT does not track with rankings on standard English-centric coding benchmarks, confirming that multilingual coding is a separate capability dimension.

Industry Insight

  • AI coding agent providers should prioritize multilingual evaluation and training beyond English, as current benchmarks significantly overestimate real-world deployment readiness in non-English markets.
  • Organizations building global software products should not assume English benchmark performance generalizes; investing in native-speaker-validated testing pipelines will be essential for multilingual agent reliability.
  • The decoupling of multilingual coding performance from general coding rankings suggests the need for specialized model architectures or fine-tuning strategies targeting locale-aware development workflows.

TL;DR

  • 提出Terminal-Bench-LILT基准,包含300个多语言编码任务,覆盖阿拉伯语、捷克语、德语、西班牙语、印地语、日语、韩语、塞尔维亚语、土耳其语和中文
  • 任务聚焦非英语软件开发特有难题,如国际化、编码处理、文本归一化和文化惯例,无直接英文对应
  • 六个前沿模型评估显示最强模型仅达63.1%通过率,大量任务未被任何模型解决
  • 多语言编码能力与通用编程基准排名无关,是独立且未被充分探索的能力维度

为什么值得看

本文揭示了当前AI编码代理在多语言环境下的严重短板,为评估和改进非英语软件开发能力提供了首个系统性基准。对AI从业者和企业而言,这标志着多语言编码能力将成为下一代AI编程工具的关键竞争维度。

技术解析

  • 数据集规模与语言覆盖:300个真实编码任务,涵盖10种语言,包括低资源语言如塞尔维亚语和捷克语
  • 任务设计原则:每个任务针对非英语软件开发特有的技术挑战,如Unicode处理、RTL文本方向、本地化日期格式、文化敏感内容等
  • 质量控制流程:所有任务由母语程序员编写,经过多阶段验证确保技术准确性和文化适当性
  • 评估结果:六个前沿模型中,最强模型仅达63.1%通过率,且不同语言间性能差异显著,与通用编程基准排名无相关性

行业启示

  • 多语言编码能力应成为AI编程代理评估的新标准维度,建议基准测试纳入多语言场景
  • 当前模型在非英语软件开发场景存在系统性缺陷,企业部署AI编码工具时需考虑本地化适配
  • 低资源语言(如塞尔维亚语、捷克语)的编码能力尤为薄弱,值得研究和资源投入

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Benchmark 基准测试 Evaluation 评测 Code Generation 代码生成 Agent Agent Dataset 数据集