Import AI 469: Science AI; RSI simulator; and Zuck's technological pessimism
DiG-bench is a new benchmark of 70 handcrafted, text-based games designed to measure AI systems' ability to autonomously discover hidden rules and mechanics in novel environments, serving as a proxy for creative intuition Frontier models (Opus 5, Fable 5 with Claude Code, GPT-5.5) show modest performance on DiG-bench, with only the top models solving Tier 7 tasks at ~20% success, far below human capability An RSI Simulator game from Paradigm Research offers an interactive way to understand the s
Analysis
TL;DR
- DiG-bench is a new benchmark of 70 handcrafted, text-based games designed to measure AI systems' ability to autonomously discover hidden rules and mechanics in novel environments, serving as a proxy for creative intuition
- Frontier models (Opus 5, Fable 5 with Claude Code, GPT-5.5) show modest performance on DiG-bench, with only the top models solving Tier 7 tasks at ~20% success, far below human capability
- An RSI Simulator game from Paradigm Research offers an interactive way to understand the strategic tradeoffs involved in recursive self-improvement and AI lab management
- Research on "inherent post-training" suggests AI systems are beginning to exhibit early signs of scientific taste, supervising frontier models in research workflows
Why It Matters
This benchmark directly addresses one of the most critical open questions in AI: whether systems can genuinely discover and generalize from novel, undocumented environments rather than relying on pre-specified rules. For researchers and practitioners, DiG-bench provides a controlled, measurable way to track progress toward autonomous discovery—a prerequisite for creativity and potentially recursive self-improvement. The article also projects human parity on DiG-bench by mid-2027, which could signal the onset of serious recursive self-improvement cycles.
Technical Details
- DiG-bench consists of 70 handcrafted, purely text-based games split into 7 difficulty tiers, with 21 publicly available and the rest held private to prevent training contamination; action spaces range from 2 to 34 options per step
- Games are designed so that both rules and objectives are hidden, requiring players to learn through exploration and interaction—similar in spirit to the visual ARC benchmark but native to language models
- Top performers: Opus 5 and Fable 5 (with Claude Code harness) lead overall, with GPT-5.5 and Kimi K3 solving select Tier 6 tasks; GLM-5.2 and Gemini 3.1 Pro manage Tier 4 levels
- The benchmark includes an optional experimentation mode allowing unrestricted interaction to help agents discover mechanics before attempting scored runs
- Research institutions involved include Oxford, Princeton, KAUST, Swiss AI Lab, Inria, and MIT, with Juergen Schmidhuber as a contributing author
Industry Insight
- DiG-bench establishes a new evaluation paradigm focused on discovery and rule inference rather than pattern recognition or retrieval, signaling a shift in how AI capabilities will be assessed as models become more capable
- The projected mid-2027 human parity timeline suggests that organizations should begin preparing for potential recursive self-improvement scenarios, including safety governance and alignment research investment
- The emergence of "scientific taste" in AI systems through inherent post-training represents an early but significant step toward autonomous AI scientists, warranting close monitoring by labs investing in AI-assisted research pipelines
Disclaimer: The above content is generated by AI and is for reference only.