Research Papers 论文研究 5h ago Updated 39m ago 更新于 39分钟前 45

Which India Survives Translation? Narrative Homogenisation Across Indian Oral Traditions in LLMs 哪种印度能在翻译中幸存?LLM中的印度口头传统叙事同质化

LLMs trained on English-dominated internet text tend to flatten diverse Indian storytelling traditions into a homogenized archetype, with cross-tradition similarity (0.52–0.66) far exceeding what genuine cultural distance would predict Prompting in regional languages (Hindi, Tamil, Bengali) consistently reduced fidelity to authentic traditions compared to English prompting, by up to 27 percentage points for Rajasthani and Bengali traditions The study examined three distinct Indian traditions—Raj LLM训练数据以英语互联网文本为主,过度代表特定文化叙事,导致非西方口头传统被扁平化为单一同质化原型 研究考察了三个印度区域传统(拉贾斯坦Pabuji史诗、泰米尔桑加姆诗歌、孟加拉民间故事)在Claude Sonnet和Gemini中的表现 跨传统相似度高达0.52-0.66,显著高于传统间真实距离预期,表明存在部分同质化现象 区域语言提示(印地语/泰米尔语/孟加拉语)反而降低了对真实传统的保真度,比英语提示低达27个百分点 该研究作为轻量级、可扩展的计算方法,补充了近期大规模人工标注的印度文化误表征研究

58
Hot 热度
72
Quality 质量
65
Impact 影响力

Analysis 深度分析

TL;DR

  • LLMs trained on English-dominated internet text tend to flatten diverse Indian storytelling traditions into a homogenized archetype, with cross-tradition similarity (0.52–0.66) far exceeding what genuine cultural distance would predict
  • Prompting in regional languages (Hindi, Tamil, Bengali) consistently reduced fidelity to authentic traditions compared to English prompting, by up to 27 percentage points for Rajasthani and Bengali traditions
  • The study examined three distinct Indian traditions—Rajasthani Pabuji epic, classical Tamil Sangam poetry, and Bengali folk tales—using Sentence-BERT embeddings and cosine similarity to measure reference drift and cross-tradition convergence
  • This finding challenges assumptions that multilingual prompting inherently improves cultural authenticity, suggesting instead that it may elicit general cultural diversity rather than simulate narrow, lesser-documented traditions
  • The pilot study offers a lightweight, scalable methodology complementary to large-scale human-annotation efforts for detecting cultural misrepresentation in LLM-generated content

Why It Matters

This research directly addresses a critical gap in AI safety and cultural representation: as LLMs are deployed globally, they risk erasing the nuance of non-Western oral and literary traditions by collapsing them into a single homogenized narrative. For practitioners building multilingual or culturally-aware systems, these findings suggest that simply switching to a regional language prompt does not guarantee authentic cultural output and may actively degrade fidelity to source traditions.

Technical Details

  • Corpus: Authentic reference corpora collected for three Indian traditions—Rajasthani Pabuji epic (11 passages), classical Tamil Sangam poetry (21 passages), and Bengali folk tales (10 passages)
  • Models tested: Claude Sonnet and Gemini, prompted with 54 generation requests across three prompt types per tradition (generic, culturally specific, and regional-language)
  • Methodology: Sentence-BERT embeddings combined with cosine similarity to compute two metrics: reference drift (output proximity to its own tradition's authentic texts vs. others) and cross-tradition convergence (similarity of outputs across traditions)
  • Key quantitative finding: Cross-tradition similarity ranged from 0.52 to 0.66, substantially higher than the distance between authentic traditions would predict; regional-language prompting reduced fidelity by up to 27 percentage points compared to English prompting for Rajasthani and Bengali traditions
  • Positioning: Framed as a lightweight, scalable complement to large-scale human-annotation studies, part of a broader doctoral research program on Indian cultural misrepresentation in LLMs

Industry Insight

  • Organizations deploying LLMs for culturally-specific content generation should not assume multilingual prompting automatically improves authenticity; rigorous evaluation against ground-truth cultural corpora is essential before relying on regional-language outputs
  • The counterintuitive finding that regional-language prompting degrades fidelity suggests current LLMs may conflate broad cultural tropes with tradition-specific nuance when prompted in non-English languages—a risk worth monitoring as multilingual capabilities expand
  • This pilot methodology (embedding-based drift and convergence measurement) offers a scalable template for auditing cultural representation across other underrepresented traditions, enabling proactive rather than reactive cultural safety assessments

TL;DR

  • LLM训练数据以英语互联网文本为主,过度代表特定文化叙事,导致非西方口头传统被扁平化为单一同质化原型
  • 研究考察了三个印度区域传统(拉贾斯坦Pabuji史诗、泰米尔桑加姆诗歌、孟加拉民间故事)在Claude Sonnet和Gemini中的表现
  • 跨传统相似度高达0.52-0.66,显著高于传统间真实距离预期,表明存在部分同质化现象
  • 区域语言提示(印地语/泰米尔语/孟加拉语)反而降低了对真实传统的保真度,比英语提示低达27个百分点
  • 该研究作为轻量级、可扩展的计算方法,补充了近期大规模人工标注的印度文化误表征研究

为什么值得看

这篇论文揭示了LLM在多语言文化叙事生成中的系统性偏差问题,对AI从业者和语言模型开发者具有重要参考价值。研究结果挑战了"多语言提示能提升文化保真度"的直觉假设,为构建更公平、更多元的AI系统提供了实证依据。

技术解析

  • 研究设计:收集了三个印度区域口头传统的真实参考语料库(11、21、10段文本),使用Claude Sonnet和Gemini两个LLM进行54次生成请求,涵盖通用、文化特定和区域语言三种提示类型
  • 评估方法:采用Sentence-BERT嵌入和余弦相似度,测量参考漂移(输出与自身传统真实文本的接近程度)和跨传统收敛(不同传统输出间的相似度)
  • 关键发现:虽然输出仍更接近自身传统参考,但跨传统相似度(0.52-0.66)远高于传统间真实距离预期,表明存在部分同质化
  • 意外结果:区域语言提示反而降低了文化保真度,英语提示表现更好,作者认为这反映了"激发一般文化多样性"与"模拟特定小众口头传统"之间的差异
  • 研究定位:作为轻量级、可扩展的计算研究,补充近期大规模人工标注的印度文化误表征研究,属于更广泛的博士研究项目的一部分

行业启示

  • LLM训练数据的语言和文化偏见问题需要系统性解决,不能仅依赖多语言提示来改善文化保真度
  • 在评估LLM的文化表现时,应区分"泛化文化多样性"与"特定小众传统模拟"的不同挑战
  • 轻量级计算研究方法可作为大规模人工标注的有效补充,为文化代表性评估提供可扩展的解决方案

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

LLM 大模型 Research 科学研究 Ethics 伦理 Dataset 数据集 Evaluation 评测