On the Navier–Stokes Millennium Prize Problem
OpenAI claims to have resolved the Navier–Stokes existence and smoothness problem, one of the seven Millennium Prize Problems, using an unreleased internal model and AI agents The achievement took approximately 88 hours of agent reasoning plus 17 hours of Lean formalization/verification via GPT-6 Astra, consuming ~130 billion output tokens for the Navier–Stokes problem alone and ~300 billion tokens across all attempted problems Tristan Buckmaster (NYU) and Levent Alpöge (Anthropic) accuse OpenAI
Analysis
TL;DR
- OpenAI claims to have resolved the Navier–Stokes existence and smoothness problem, one of the seven Millennium Prize Problems, using an unreleased internal model and AI agents
- The achievement took approximately 88 hours of agent reasoning plus 17 hours of Lean formalization/verification via GPT-6 Astra, consuming ~130 billion output tokens for the Navier–Stokes problem alone and ~300 billion tokens across all attempted problems
- Tristan Buckmaster (NYU) and Levent Alpöge (Anthropic) accuse OpenAI of scooping work they had been developing for nearly a year using Claude and Codex (GPT-5.6 Sol), raising concerns about data usage and competitive ethics
- OpenAI denies accessing any specific user data but acknowledges it "cannot rule out" that de-identified data from users' interactions with their products may have improved model performance
- The incident draws a parallel to computer security, where mere rumors of vulnerabilities can trigger expensive agent-driven exploit searches, suggesting a new paradigm where mathematical breakthroughs may be preempted by AI racing
Why It Matters
This event represents a potential watershed moment in AI-assisted mathematical research, demonstrating that large-scale AI agent systems can tackle problems at the frontier of human mathematical knowledge. It also exposes critical ethical and legal ambiguities around how AI labs use user-generated data and interactions, with direct implications for researchers, institutions, and the broader scientific community.
Technical Details
- OpenAI deployed AI agents that sent 4.9 million messages across all attempted Millennium Prize problems, with the Navier–Stokes resolution alone consuming 2.7 million messages and approximately 130 billion output tokens
- Lean formalization and proof verification was handled by GPT-6 Astra over an additional 17 hours, indicating a two-stage pipeline: agent-based reasoning followed by formal verification
- The total computational expenditure is estimated at roughly $15,000,000 if priced at public API rates for GPT-6 Astra, though the actual cost of the unreleased internal model remains undisclosed
- Buckmaster and Alpöge's approach relied extensively on Claude and Codex (primarily GPT-5.6 Sol) over nearly a year of collaborative work, contrasting with OpenAI's rapid 88-hour sprint
- OpenAI notes their proofs differ significantly from Buckmaster and Alpöge's, including in the Euler case (forced vs. unforced), suggesting independent derivation paths
Industry Insight
- AI labs should establish transparent data usage policies and opt-in frameworks for researchers using their platforms on high-stakes problems, as the current ambiguity around "de-identified data improving models" creates reputational and legal risk
- The "rumor-driven breakthrough" dynamic mirrors emerging patterns in cybersecurity and may become standard in mathematics, incentivizing labs to monitor academic discourse and rumors as intelligence signals for competitive AI research deployment
- Institutions and funding bodies should develop guidelines for AI-assisted mathematical authorship and priority claims, as traditional norms around discovery, collaboration, and attribution are ill-equipped for scenarios where AI agents can reproduce or preempt human-led research at scale
Disclaimer: The above content is generated by AI and is for reference only.