Research Papers 论文研究 4h ago Updated 31m ago 更新于 31分钟前 48

The Limits of Automatic Evaluation of Creativity in Large Language Models 大语言模型中创造力自动评估的局限性

LLM-based judges show a systematic bias favoring AI-generated stories over human-authored ones, prioritizing stylistic polish over unpredictability and creative novelty Widely used automatic evaluation metrics exhibit near-zero correlation with human creativity judgments across both human- and AI-generated texts Human evaluations were collected across 11 distinct dimensions of creativity using the WritingPrompts dataset, revealing the multidimensional nature of creative assessment The study demo 研究揭示当前自动评估方法与人类创造力判断存在显著偏差,LLM-as-a-Judge系统性偏好AI生成文本 广泛使用的自动指标与人类判断几乎零相关性,无法捕捉创造力的重要维度 LLM judges倾向于AI文本的风格特征,而非人类文本的不可预测性等品质 创造力的多维性和主观性难以简化为单一计算指标

65
Hot 热度
75
Quality 质量
68
Impact 影响力

Analysis 深度分析

TL;DR

  • LLM-based judges show a systematic bias favoring AI-generated stories over human-authored ones, prioritizing stylistic polish over unpredictability and creative novelty
  • Widely used automatic evaluation metrics exhibit near-zero correlation with human creativity judgments across both human- and AI-generated texts
  • Human evaluations were collected across 11 distinct dimensions of creativity using the WritingPrompts dataset, revealing the multidimensional nature of creative assessment
  • The study demonstrates fundamental limitations in reducing creativity to computational metrics, highlighting a significant gap between automated and human evaluation paradigms

Why It Matters

This research directly challenges the growing reliance on LLM-as-a-Judge frameworks and automated metrics for evaluating creative AI outputs, which are increasingly used in benchmarking and product development. For AI practitioners, it signals that current evaluation pipelines may be producing misleadingly positive assessments of creative capabilities, potentially driving optimization toward superficial stylistic features rather than genuine creative quality.

Technical Details

  • Dataset: Human- and AI-generated short stories from the WritingPrompts dataset, evaluated across 11 dimensions of creativity by human raters
  • Evaluation methods compared: (1) Human evaluations, (2) LLM-as-a-Judge automated evaluations, (3) Objective automatic metrics (e.g., perplexity, BLEU, ROUGE, or similar standard NLP metrics)
  • Key finding on LLM judges: Systematic preference for AI-generated content, consistently favoring stylistic characteristics such as fluency and coherence over human-text qualities like unpredictability and originality
  • Correlation analysis: Near-zero alignment between automatic metrics and human judgments across both human- and AI-generated stories, indicating these metrics fail to capture essential creativity dimensions
  • Subject area: Computation and Language (cs.CL), Artificial Intelligence (cs.AI), Computers and Society (cs.CY)

Industry Insight

  • Organizations relying on automated benchmarks to measure creative AI capabilities should incorporate human-in-the-loop evaluation, particularly for dimensions like unpredictability and novelty that current metrics miss
  • The bias in LLM-as-a-Judge systems toward AI-generated text suggests a feedback loop risk: models may be optimized against evaluators that inherently favor their own output patterns, potentially stalling genuine creative advancement
  • Researchers and product teams should develop or adopt creativity-specific evaluation frameworks that explicitly measure the 11+ dimensions humans use, rather than relying on generic text quality metrics

TL;DR

  • 研究揭示当前自动评估方法与人类创造力判断存在显著偏差,LLM-as-a-Judge系统性偏好AI生成文本
  • 广泛使用的自动指标与人类判断几乎零相关性,无法捕捉创造力的重要维度
  • LLM judges倾向于AI文本的风格特征,而非人类文本的不可预测性等品质
  • 创造力的多维性和主观性难以简化为单一计算指标

为什么值得看

这篇论文直接挑战了当前AI社区广泛依赖的LLM-as-a-Judge评估范式,揭示了其在创造力评估方面的系统性偏差。对于AI从业者和研究者而言,这提醒我们在评估模型创造力表现时需要谨慎对待自动指标,避免被有偏的评估结果误导。

技术解析

  • 研究使用WritingPrompts数据集,收集人类和AI生成的短篇故事,在11个创造力维度上进行人类评估
  • 对比三种评估方式:人类评估、自动化客观指标、LLM-as-a-Judge,发现自动评估与人类判断存在显著偏差
  • LLM judges表现出系统性偏见, consistently favor AI-generated stories over human-authored texts
  • 相关性分析显示,常用自动指标与人类创造力判断几乎零对齐(near-zero alignment)

行业启示

  • 当前流行的LLM-as-a-Judge评估方法在创造力维度存在系统性偏差,需谨慎使用或进行偏差校正
  • 自动评估指标在衡量创造力等主观性强的维度上存在根本性局限,建议结合人类评估或多维度评估框架
  • 研究提示AI社区需要开发更贴近人类创造力认知的评估标准,而非简单依赖现有自动化方案

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

LLM 大模型 Evaluation 评测 Creative AI 创意AI Research 科学研究 Dataset 数据集