Research Papers 论文研究 5h ago Updated 37m ago 更新于 37分钟前 47

Evaluating Language Models in Realistic Conversational Contexts 评估语言模型在真实对话上下文中的表现

UPHELD is introduced as a large, reference-full benchmark for evaluating human-scale conversational ability beyond factual correctness The dataset contains hundreds of complete human-to-human dialogues authored by professional script writers with 36,000+ per-turn human annotations across 30,000+ expert-generated dialogue turns Classical automatic metrics and reference-free LLM-as-a-judge approaches were found unreliable when correlated with expert human judgment A Mixture-of-Judges framework was 提出UPHELD基准,包含30,000+专家生成对话轮次和36,000+人工标注,用于评估人类规模对话能力 现有自动指标和LLM-as-a-judge方法与人类判断相关性不可靠 开发Mixture-of-Judges框架,结合多种评估信号,将相关性提升约30%

62
Hot 热度
72
Quality 质量
68
Impact 影响力

Analysis 深度分析

TL;DR

  • UPHELD is introduced as a large, reference-full benchmark for evaluating human-scale conversational ability beyond factual correctness
  • The dataset contains hundreds of complete human-to-human dialogues authored by professional script writers with 36,000+ per-turn human annotations across 30,000+ expert-generated dialogue turns
  • Classical automatic metrics and reference-free LLM-as-a-judge approaches were found unreliable when correlated with expert human judgment
  • A Mixture-of-Judges framework was developed that combines multiple evaluative signals and improves correlation with human assessments by approximately 30%
  • UPHELD fills a crucial gap in the LLM dataset landscape by providing a robust, human-grounded foundation for evaluating conversational intelligence

Why It Matters

As LLMs are increasingly deployed for open-ended, multi-turn interactions, the lack of reliable evaluation frameworks for human-scale dialogue quality represents a significant bottleneck. This work directly addresses that gap by providing both a high-quality benchmark and improved evaluation methodologies, enabling practitioners to better assess conversational performance before deployment.

Technical Details

  • UPHELD (UPwork Human-Scale Evaluated Long Dialogues) consists of hundreds of complete human-to-human dialogues authored by professional script writers, featuring realistic turn densities
  • The benchmark includes 36,000+ per-turn human annotations across 30,000+ expert-generated dialogue turns, providing dense human-grounded evaluation signals
  • Systematic evaluation of classical automatic metrics and reference-free LLM-as-a-judge approaches revealed poor correlation with expert human judgment
  • The Mixture-of-Judges framework combines multiple evaluative signals to improve correlation with human assessments by approximately 30%

Industry Insight

  • The unreliability of existing LLM-as-a-judge approaches for conversational evaluation suggests practitioners should adopt ensemble or mixture-based evaluation strategies rather than relying on single-judge systems
  • The 30% improvement from the Mixture-of-Judges framework indicates that combining diverse evaluative signals is critical for accurate conversational quality assessment
  • Organizations deploying conversational LLMs should prioritize human-annotated benchmarks like UPHELD for validation, as synthetic evaluation sources may produce misleading quality signals

TL;DR

  • 提出UPHELD基准,包含30,000+专家生成对话轮次和36,000+人工标注,用于评估人类规模对话能力
  • 现有自动指标和LLM-as-a-judge方法与人类判断相关性不可靠
  • 开发Mixture-of-Judges框架,结合多种评估信号,将相关性提升约30%

为什么值得看

本文填补了LLM对话评估领域的关键空白,提供了首个大规模、人类标注的对话质量基准。对开发对话型AI系统的团队具有重要参考价值,有助于建立更可靠的评估体系。

技术解析

  • UPHELD(UPwork Human-Scale Evaluated Long Dialogues)基准由专业编剧撰写数百个完整的人与人对话,具有真实的对话轮次密度
  • 包含36,000+每轮人工标注和30,000+专家生成的对话轮次,覆盖多轮对话的一致性评估
  • 系统评估了经典自动指标和参考自由的LLM-as-a-judge方法,发现其在人类判断相关性上不可靠
  • 开发Mixture-of-Judges框架,整合多种评估信号,将相关性提升约30%

行业启示

  • 对话评估需从合成数据转向真实人类标注,现有基准存在根本性缺陷
  • 单一评估方法存在局限,多信号融合是提升评估可靠性的有效路径
  • 为LLM对话能力评估提供了新的基准和方法论,推动行业向更科学的评估体系演进

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Conversational AI 对话系统 Evaluation 评测 LLM 大模型 Benchmark 基准测试 Research 科学研究