New math benchmark reveals AI models confidently solve problems that have no solution
A team of 64 mathematicians constructed a new AI benchmark called SOOHAK, comprising 439 handwritten mathematical tasks, of which 99 were intentionally designed to be unsolvable. This test aims to evaluate AI models not only in solving problems but also in recognizing whether the problem itself is valid. Currently, Google's Gemini 3 Pro leads in research-level problems but achieves only a 30% accuracy rate. More notably, in identifying unsolvable tasks, no model surpasses 50% accuracy. The research found that increasing computational resources enhances models' problem-solving abilities but does not improve their capability to identify unsolvable problems. The purpose of the SOOHAK benchmark is to explicitly quantify the significant gap that exists in current AI systems between sporadic flashes of brilliance and comprehensive mastery of research skills.
Analysis
Sixty-four mathematicians hand-penned 439 problems to build SOOHAK, a new benchmark designed to humble AI. Their masterstroke wasn’t the hardest problems—it was embedding 99 that are deliberately unsolvable. The results are a damning portrait of artificial intelligence today: it’s becoming a brilliant parrot that can solve research-level questions but cannot fathom the concept of a trick question. This isn’t just a gap in knowledge; it’s a chasm in judgment, revealing a fundamental flaw in how we’re building and evaluating these systems.
Google’s Gemini 3 Pro, leading the pack, solves a respectable 30% of the research-level tasks. That’s a genuinely impressive feat, a testament to the raw computational horsepower and pattern-matching prowess we’ve achieved. It can navigate the dense terrain of advanced mathematics with growing skill. Yet, the sobering statistic is that no model cracks the 50% mark when it comes to spotting problems that have no answer. More compute, the usual sledgehammer for AI progress, makes them better at solving the solvable. It does nothing to improve their humility in the face of the impossible. This is the core paradox: we are building systems that are increasingly capable and simultaneously increasingly confident in their own potential for error.
Think about what this means in practice. An AI that can prove a theorem but cannot identify a flawed premise is not a research assistant; it’s a very sophisticated, very confident idiot. It operates on the assumption that every problem handed to it has a neat, computable solution. This is the "optimizer’s curse" baked into the neural architecture. These models are trained on vast oceans of human-generated answers, not on a rich understanding of why some questions are bad. They learn the syntax of reasoning but miss the semantics of doubt. SOOHAK brilliantly isolates this. The benchmark isn’t testing intelligence; it’s testing for the presence of wisdom, and the results show the tank is running on empty.
The implications stretch far beyond mathematics. Consider the burgeoning field of AI for science. A model tasked with drug discovery might generate thousands of promising molecular structures. But can it tell you when a hypothesis is fundamentally incoherent? Can it recognize an experimental setup that is logically designed to yield no meaningful data? If it cannot distinguish a solvable math problem from a nonsensical one, we have zero reason to trust its judgment in ambiguous, real-world domains where the "ground truth" is murkier than a proof. We’re building a generation of tools that are expert at finding answers but incompetent at identifying questions that shouldn’t be asked.
This exposes a deep flaw in our current philosophy of AI evaluation. We obsess over leaderboards for MMLU, GSM8K, and now SOOHAK’s solvable tier. We reward systems for getting the right answer faster. But we have neglected to build robust, standardized gauges for a model’s epistemic humility—its ability to say, “I don’t know,” or more precisely, “This question is nonsense.” The industry’s mantra is to scale compute and data, but SOOHAK suggests a different scaling challenge: scaling judgment. How do you train a model to recognize the void? It’s not a matter of more parameters; it’s a matter of a different kind of training signal entirely, perhaps one focused on error analysis, adversarial validation, and the detection of logical inconsistencies rather than their resolution.
The 64 mathematicians didn’t just create a test; they held up a mirror. The reflection shows an AI landscape that is brilliant at performing competence but has not yet learned the first lesson of real expertise: knowing the boundaries of the problem space. The confidence these models exude is a trained artifact, not an earned one. Every time a model selects an answer for an unsolvable problem, it’s not just wrong—it’s exhibiting a profound misunderstanding of the nature of inquiry itself.
This will be the next great bottleneck. After we’ve optimized for accuracy and speed, we must optimize for reliability of judgment. The future of trustworthy AI isn’t just in systems that can solve the hardest problems we throw at them, but in systems that can tell us when we’ve asked a bad question. Until then, we’re left with powerful oracles that will confidently whisper answers to queries that deserve only silence. The race to build artificial general intelligence might be less about creating a godlike solver and more about instilling a very human capacity for skepticism. SOOHAK is the first serious benchmark for that, and our current models are flunking.
Disclaimer: The above content is generated by AI and is for reference only.