You Really Didn't Get That? Benchmarking Social Pragmatic Inference for Indirect and Playful Chinese Online Comments
Introduces a new benchmark for evaluating LLMs' ability to recover situated pragmatic meanings in indirect and playful Chinese online comments Dataset constructed from over 200,000 public Chinese social media interactions, yielding 4,735 human-validated diagnostic items pairing target comments with reconstructed context and plausible misreadings Eight LLMs evaluated in a cross-writer setting; strongest model achieves 81.42% leave-writer-out accuracy, with mean accuracy of 68.70% versus 90.8% for
Analysis
TL;DR
- Introduces a new benchmark for evaluating LLMs' ability to recover situated pragmatic meanings in indirect and playful Chinese online comments
- Dataset constructed from over 200,000 public Chinese social media interactions, yielding 4,735 human-validated diagnostic items pairing target comments with reconstructed context and plausible misreadings
- Eight LLMs evaluated in a cross-writer setting; strongest model achieves 81.42% leave-writer-out accuracy, with mean accuracy of 68.70% versus 90.8% for humans
- Models tend to recognize broad irony or playfulness but frequently misidentify the specific pragmatic mechanism or interactional move
- Highlights a significant gap between human and machine performance in naturalistic social pragmatic inference
Why It Matters
This benchmark addresses a critical gap in LLM evaluation by moving beyond controlled, predefined pragmatic categories to test models on naturally occurring, context-dependent social language. For AI practitioners building multilingual or culturally-aware systems, it reveals that current models still struggle with the nuanced, situated interpretation that human users handle effortlessly. The findings underscore the need for better evaluation frameworks that capture real-world communication complexity rather than artificial test scenarios.
Technical Details
- Dataset construction: 4,735 diagnostic items extracted from 200,000+ public Chinese social media interaction records, each item pairing a target comment with reconstructed preceding context and multiple plausible misreadings
- Evaluation protocol: Cross-writer setting where LLMs serve as both question writers and solvers, using leave-writer-out accuracy to prevent data leakage and author-specific bias
- Models evaluated: Eight LLMs tested, with the best achieving 81.42% and the mean across all models at 68.70%, compared to 90.8% human baseline
- Error analysis: Case studies reveal a consistent pattern where models detect surface-level irony or playfulness but fail to correctly identify the specific pragmatic mechanism (e.g., sarcasm type, teasing strategy, face-saving move) or the interactional function within the exchange
- Benchmarks and metrics: Leave-writer-out accuracy as the primary metric, designed to measure generalization across different writers and contexts rather than memorization
Industry Insight
- The 22-point gap between human (90.8%) and best model (81.42%) accuracy on pragmatic inference suggests that social language understanding remains a significant bottleneck for deploying LLMs in culturally nuanced, multilingual applications—particularly in Chinese-language contexts where indirect communication is prevalent
- The cross-writer evaluation design should become a standard practice in benchmark development to prevent overestimation of model capabilities through author-specific patterns
- Developers building chatbots, content moderation systems, or social media tools for Chinese-speaking audiences should prioritize training and evaluation on situated pragmatic tasks rather than relying solely on general language benchmarks, as these models may appear competent on standard metrics while failing in real conversational settings
Disclaimer: The above content is generated by AI and is for reference only.