AI models flub these intelligence tests. Can you fare any better?
Puzzles and games have served as central AI benchmarks since 1959, with machine learning popularized through Arthur Samuel's checkers program AI puzzle-solving capabilities have improved dramatically, with NYT Connections puzzle accuracy jumping from 18% in late 2024 to near-perfect by early 2025 Despite rapid progress, AI models still significantly underperform humans in spatial reasoning, particularly 3D mental rotation tasks Frontier LLMs struggle with memory-based puzzles where training data
Analysis
TL;DR
- Puzzles and games have served as central AI benchmarks since 1959, with machine learning popularized through Arthur Samuel's checkers program
- AI puzzle-solving capabilities have improved dramatically, with NYT Connections puzzle accuracy jumping from 18% in late 2024 to near-perfect by early 2025
- Despite rapid progress, AI models still significantly underperform humans in spatial reasoning, particularly 3D mental rotation tasks
- Frontier LLMs struggle with memory-based puzzles where training data memorization interferes with solving novel variations of familiar problems
- Visual and abstract reasoning remain major weaknesses, as demonstrated by continued failures on ARC-AGI benchmarks even as performance improves
Why It Matters
This article provides a comprehensive assessment of AI capabilities through the lens of puzzle-solving, offering practitioners a practical framework for understanding where current models excel and where fundamental gaps remain. The rapid improvement trajectory in certain puzzle domains combined with persistent failures in others reveals important insights about the nature of current AI intelligence and its limitations.
Technical Details
- Spatial Reasoning Deficits: Language models with visual capabilities still fail abysmally at mental rotation problems requiring 3D object manipulation, despite advances in world models for physical environment understanding
- Memory Interference: Research from Google and UIUC (2024) demonstrated that frontier LLMs trained on Knights and Knaves puzzle variations often rely on memorized patterns rather than logical reasoning, causing failures on subtly modified problems
- ARC-AGI Performance: Models perform better when grid puzzles are presented as numerical strings rather than visual images, but often derive non-generalizable rules rather than the simple visual concepts humans use
- Rapid Improvement Trajectory: Columbia University's late 2024 study showed 18% accuracy on NYT Connections puzzles, while early 2025 models achieved near-perfect performance, indicating exponential capability growth
- SimpleBench Vulnerabilities: Even top-tier models trip on problems resembling training data, where humans easily spot tricks that models miss due to pattern-matching over general reasoning
Industry Insight
- The dramatic improvement in puzzle-solving suggests AI capabilities are advancing faster than many practitioners expect, but persistent failures in spatial and visual reasoning indicate these areas require fundamentally different architectural approaches rather than just scale
- Memory interference problems reveal that current LLMs lack true reasoning capabilities and instead rely on pattern matching, suggesting the need for better evaluation benchmarks that can distinguish memorization from genuine understanding
- The gap between human and AI performance on abstract visual reasoning (ARC-AGI) points to opportunities for research in neuro-symbolic approaches and hybrid architectures that combine neural pattern recognition with explicit rule-based reasoning
Disclaimer: The above content is generated by AI and is for reference only.