Research Papers 论文研究 5h ago Updated 1h ago 更新于 1小时前 46

You Really Didn't Get That? Benchmarking Social Pragmatic Inference for Indirect and Playful Chinese Online Comments 你真的没懂吗?针对间接和幽默中文网络评论的社会语用推理基准测试

Introduces a new benchmark for evaluating LLMs' ability to recover situated pragmatic meanings in indirect and playful Chinese online comments Dataset constructed from over 200,000 public Chinese social media interactions, yielding 4,735 human-validated diagnostic items pairing target comments with reconstructed context and plausible misreadings Eight LLMs evaluated in a cross-writer setting; strongest model achieves 81.42% leave-writer-out accuracy, with mean accuracy of 68.70% versus 90.8% for 提出首个针对中文网络评论间接语用推理的基准测试,聚焦讽刺、幽默等隐含社会意义的理解能力评估 从20万+中文社交媒体互动记录中构建4,735个人工验证的诊断样本,每个样本包含目标评论、重建上下文和合理误读选项 在跨作者设置下评估8个LLM,最强模型准确率达81.42%,平均68.70%,人类基准为90.8% 案例分析显示模型能识别广泛讽刺或幽默,但常错误判断具体机制或互动行为类型

62
Hot 热度
72
Quality 质量
65
Impact 影响力

Analysis 深度分析

TL;DR

  • Introduces a new benchmark for evaluating LLMs' ability to recover situated pragmatic meanings in indirect and playful Chinese online comments
  • Dataset constructed from over 200,000 public Chinese social media interactions, yielding 4,735 human-validated diagnostic items pairing target comments with reconstructed context and plausible misreadings
  • Eight LLMs evaluated in a cross-writer setting; strongest model achieves 81.42% leave-writer-out accuracy, with mean accuracy of 68.70% versus 90.8% for humans
  • Models tend to recognize broad irony or playfulness but frequently misidentify the specific pragmatic mechanism or interactional move
  • Highlights a significant gap between human and machine performance in naturalistic social pragmatic inference

Why It Matters

This benchmark addresses a critical gap in LLM evaluation by moving beyond controlled, predefined pragmatic categories to test models on naturally occurring, context-dependent social language. For AI practitioners building multilingual or culturally-aware systems, it reveals that current models still struggle with the nuanced, situated interpretation that human users handle effortlessly. The findings underscore the need for better evaluation frameworks that capture real-world communication complexity rather than artificial test scenarios.

Technical Details

  • Dataset construction: 4,735 diagnostic items extracted from 200,000+ public Chinese social media interaction records, each item pairing a target comment with reconstructed preceding context and multiple plausible misreadings
  • Evaluation protocol: Cross-writer setting where LLMs serve as both question writers and solvers, using leave-writer-out accuracy to prevent data leakage and author-specific bias
  • Models evaluated: Eight LLMs tested, with the best achieving 81.42% and the mean across all models at 68.70%, compared to 90.8% human baseline
  • Error analysis: Case studies reveal a consistent pattern where models detect surface-level irony or playfulness but fail to correctly identify the specific pragmatic mechanism (e.g., sarcasm type, teasing strategy, face-saving move) or the interactional function within the exchange
  • Benchmarks and metrics: Leave-writer-out accuracy as the primary metric, designed to measure generalization across different writers and contexts rather than memorization

Industry Insight

  • The 22-point gap between human (90.8%) and best model (81.42%) accuracy on pragmatic inference suggests that social language understanding remains a significant bottleneck for deploying LLMs in culturally nuanced, multilingual applications—particularly in Chinese-language contexts where indirect communication is prevalent
  • The cross-writer evaluation design should become a standard practice in benchmark development to prevent overestimation of model capabilities through author-specific patterns
  • Developers building chatbots, content moderation systems, or social media tools for Chinese-speaking audiences should prioritize training and evaluation on situated pragmatic tasks rather than relying solely on general language benchmarks, as these models may appear competent on standard metrics while failing in real conversational settings

TL;DR

  • 提出首个针对中文网络评论间接语用推理的基准测试,聚焦讽刺、幽默等隐含社会意义的理解能力评估
  • 从20万+中文社交媒体互动记录中构建4,735个人工验证的诊断样本,每个样本包含目标评论、重建上下文和合理误读选项
  • 在跨作者设置下评估8个LLM,最强模型准确率达81.42%,平均68.70%,人类基准为90.8%
  • 案例分析显示模型能识别广泛讽刺或幽默,但常错误判断具体机制或互动行为类型

为什么值得看

该研究填补了中文网络语用推理评估的空白,揭示了当前LLM在理解隐含社会意义方面的显著差距。对AI从业者而言,这为评估模型真实社交理解能力提供了可量化的诊断工具,而非仅依赖传统NLP基准。

技术解析

  • 数据集构建:从200,000+公开中文社交媒体互动记录中筛选并构建4,735个人工验证的诊断项,每个样本包含目标评论、重建的前置上下文和多个合理误读选项
  • 评估范式:采用跨作者设置(leave-writer-out),将LLM同时作为问题编写者和解题者,测试模型能否区分自然评论在特定对话中的实际语用功能
  • 模型表现:8个主流LLM参与评估,最强模型达到81.42%准确率,整体平均68.70%,与人类90.8%准确率存在显著差距
  • 错误模式分析:模型普遍能识别评论的宏观语用特征(如讽刺、幽默),但在具体机制识别(如反讽、双关、调侃)和互动行为分类上频繁出错

行业启示

  • 语用能力成为新瓶颈:当LLM在显式语言任务上接近人类水平后,隐含社会意义的理解将成为下一代模型竞争的关键维度,建议将语用推理纳入模型评估体系
  • 中文网络语言研究价值凸显:中文特有的间接表达、网络梗文化和语境依赖性强,该基准为中文AI能力评估提供了本土化、高价值的测试场景
  • 人机差距仍有优化空间:68.70% vs 90.8%的准确率差距表明当前模型在真实社交场景中的理解能力有限,需加强上下文建模和语用知识注入

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Evaluation 评测 Benchmark 基准测试 Dataset 数据集 Conversational AI 对话系统 Research 科学研究