The Limits of Automatic Evaluation of Creativity in Large Language Models
LLM-based judges show a systematic bias favoring AI-generated stories over human-authored ones, prioritizing stylistic polish over unpredictability and creative novelty Widely used automatic evaluation metrics exhibit near-zero correlation with human creativity judgments across both human- and AI-generated texts Human evaluations were collected across 11 distinct dimensions of creativity using the WritingPrompts dataset, revealing the multidimensional nature of creative assessment The study demo
Analysis
TL;DR
- LLM-based judges show a systematic bias favoring AI-generated stories over human-authored ones, prioritizing stylistic polish over unpredictability and creative novelty
- Widely used automatic evaluation metrics exhibit near-zero correlation with human creativity judgments across both human- and AI-generated texts
- Human evaluations were collected across 11 distinct dimensions of creativity using the WritingPrompts dataset, revealing the multidimensional nature of creative assessment
- The study demonstrates fundamental limitations in reducing creativity to computational metrics, highlighting a significant gap between automated and human evaluation paradigms
Why It Matters
This research directly challenges the growing reliance on LLM-as-a-Judge frameworks and automated metrics for evaluating creative AI outputs, which are increasingly used in benchmarking and product development. For AI practitioners, it signals that current evaluation pipelines may be producing misleadingly positive assessments of creative capabilities, potentially driving optimization toward superficial stylistic features rather than genuine creative quality.
Technical Details
- Dataset: Human- and AI-generated short stories from the WritingPrompts dataset, evaluated across 11 dimensions of creativity by human raters
- Evaluation methods compared: (1) Human evaluations, (2) LLM-as-a-Judge automated evaluations, (3) Objective automatic metrics (e.g., perplexity, BLEU, ROUGE, or similar standard NLP metrics)
- Key finding on LLM judges: Systematic preference for AI-generated content, consistently favoring stylistic characteristics such as fluency and coherence over human-text qualities like unpredictability and originality
- Correlation analysis: Near-zero alignment between automatic metrics and human judgments across both human- and AI-generated stories, indicating these metrics fail to capture essential creativity dimensions
- Subject area: Computation and Language (cs.CL), Artificial Intelligence (cs.AI), Computers and Society (cs.CY)
Industry Insight
- Organizations relying on automated benchmarks to measure creative AI capabilities should incorporate human-in-the-loop evaluation, particularly for dimensions like unpredictability and novelty that current metrics miss
- The bias in LLM-as-a-Judge systems toward AI-generated text suggests a feedback loop risk: models may be optimized against evaluators that inherently favor their own output patterns, potentially stalling genuine creative advancement
- Researchers and product teams should develop or adopt creativity-specific evaluation frameworks that explicitly measure the 11+ dimensions humans use, rather than relying on generic text quality metrics
Disclaimer: The above content is generated by AI and is for reference only.