Are LLMs becoming similarly creative? Evidence from three years of models
LLM creative outputs show a statistically significant decrease in diversity over a three-year period, suggesting convergence in creative substance across models The study uses sentence-embedding similarity to analyze responses from Infinity-Chat100 (real-world open-ended queries) and the Alternate Uses Task (psychometric creativity assessment) Unlike traditional benchmarks focused on verifiable answers, this research evaluates open-ended tasks where originality and diversity are as critical as q
Analysis
TL;DR
- LLM creative outputs show a statistically significant decrease in diversity over a three-year period, suggesting convergence in creative substance across models
- The study uses sentence-embedding similarity to analyze responses from Infinity-Chat100 (real-world open-ended queries) and the Alternate Uses Task (psychometric creativity assessment)
- Unlike traditional benchmarks focused on verifiable answers, this research evaluates open-ended tasks where originality and diversity are as critical as quality
- The findings raise concerns about LLM-driven homogenization potentially diminishing human agency in human-AI co-creative workflows
- This represents a preliminary but important shift in how we evaluate LLM capabilities beyond factual accuracy
Why It Matters
As LLMs become increasingly integrated into creative and ideation workflows, understanding whether these systems are converging toward similar outputs has direct implications for their utility in human creative processes. For AI practitioners and researchers, this finding challenges the assumption that newer models are inherently more capable across all dimensions, particularly in open-ended domains where diversity and originality are paramount. The trend toward homogenization could limit the value of LLMs as creative partners rather than mere content generators.
Technical Details
- Datasets: Infinity-Chat100, a real-world collection of open-ended user queries, and the Alternate Uses Task, a well-established psychometric creativity assessment from psychology
- Methodology: Sentence-embedding similarity analysis was used to measure diversity and convergence in model outputs across three years of LLM releases
- Scope: The study spans three years of model releases, tracking trends rather than providing a single-point comparison
- Evaluation focus: Open-ended tasks where creativity, originality, and diversity matter as much as quality, contrasting with traditional benchmarks that prioritize verifiable answers
- Statistical significance: The decrease in output diversity over time was found to be statistically significant, strengthening the validity of the convergence claim
Industry Insight
- AI developers should consider diversity-preserving techniques (e.g., controlled temperature variation, diversity-promoting fine-tuning, or ensemble approaches) to counteract homogenization trends in creative applications
- Organizations relying on LLMs for ideation and co-creative work should audit model outputs for convergence and potentially maintain access to older or more diverse models as part of their creative toolkit
- The research community needs new evaluation frameworks that go beyond accuracy metrics to capture creativity, originality, and output diversity, especially as LLMs are deployed more heavily in human-AI collaborative creative workflows
Disclaimer: The above content is generated by AI and is for reference only.