Artificial Analysis overhauls its Intelligence Index after GPT-6 Astra scoring drew skepticism
Artificial Analysis released Intelligence Index v4.2, updating GPT-6 Astra's score to reflect a four-point gain over its predecessor, contradicting earlier assessments that placed it on par The overhaul was prompted by widespread criticism that the previous benchmark failed to capture Astra's actual capabilities, while other evaluations (Epoch AI, ARC-AGI-3) had ranked it far ahead Two new benchmarks were added (AA-Briefcase for real-world knowledge work and GDP.pdf from Surge AI for PDF analysi
Analysis
TL;DR
- Artificial Analysis released Intelligence Index v4.2, updating GPT-6 Astra's score to reflect a four-point gain over its predecessor, contradicting earlier assessments that placed it on par
- The overhaul was prompted by widespread criticism that the previous benchmark failed to capture Astra's actual capabilities, while other evaluations (Epoch AI, ARC-AGI-3) had ranked it far ahead
- Two new benchmarks were added (AA-Briefcase for real-world knowledge work and GDP.pdf from Surge AI for PDF analysis), while GPQA-Diamond was dropped as models had solved it
- Private test data now accounts for 40% of weighting to reduce benchmark gaming, and scoring errors across multiple benchmarks were corrected
- Claude Fable 5.1 retains the top ranking, with Astra in second and Meta in third; cost-to-performance leaders include Anthropic, OpenAI, Meta, and Zhipu AI
Why It Matters
This update highlights the growing tension between benchmark design and real model capability assessment as frontier models advance rapidly. For AI practitioners, it underscores the importance of using multiple evaluation sources rather than relying on a single leaderboard, and signals that benchmark methodologies must evolve continuously to stay relevant. The shift toward private test data and real-world task benchmarks reflects an industry-wide push to produce more meaningful performance measurements.
Technical Details
- Benchmark additions and removals: AA-Briefcase evaluates real-world knowledge work performance, while GDP.pdf (from Surge AI) tests PDF document analysis capabilities; GPQA-Diamond was removed due to model saturation
- Scoring methodology changes: Private test data now comprises 40% of the overall weighting to make benchmark gaming more difficult, and grading systems were tweaked across multiple benchmarks for more stable results
- Ranking outcomes: Claude Fable 5.1 leads the index, GPT-6 Astra ranks second with a four-point improvement over predecessor Sol, and Meta holds third place
- Cost-efficiency findings: Anthropic, OpenAI, Meta, and Zhipu AI share the lead on cost-to-performance ratio, while Astra uses fewer tokens per task than all other frontier models
- Version 5 development: A major overhaul has been in progress for eight months and will be released in stages, indicating ongoing methodological refinement
Industry Insight
- Benchmark providers face increasing pressure to adapt quickly as model capabilities outpace evaluation cycles; organizations should monitor multiple leaderboards rather than relying on a single source of truth
- The inclusion of real-world task benchmarks (AA-Briefcase, GDP.pdf) signals a broader industry shift toward practical, application-oriented evaluation over purely academic or trivia-style tests
- The 40% private test data weighting suggests benchmark gaming remains a significant concern, and practitioners should be cautious about models that appear to overperform on public leaderboards without independent verification
Disclaimer: The above content is generated by AI and is for reference only.