AI News AI资讯 3mo ago Updated 2mo ago 更新于 2个月前 85

New math benchmark reveals AI models confidently solve problems that have no solution 新数学基准测试揭示AI模型自信解答无解问题

A team of 64 mathematicians constructed a new AI benchmark called SOOHAK, comprising 439 handwritten mathematical tasks, of which 99 were intentionally designed to be unsolvable. This test aims to evaluate AI models not only in solving problems but also in recognizing whether the problem itself is valid. Currently, Google's Gemini 3 Pro leads in research-level problems but achieves only a 30% accuracy rate. More notably, in identifying unsolvable tasks, no model surpasses 50% accuracy. The research found that increasing computational resources enhances models' problem-solving abilities but does not improve their capability to identify unsolvable problems. The purpose of the SOOHAK benchmark is to explicitly quantify the significant gap that exists in current AI systems between sporadic flashes of brilliance and comprehensive mastery of research skills. 由64位数学家组成的团队构建了名为SOOHAK的新型AI基准测试,包含439个手写数学任务,其中99个被故意设计为无解问题。该测试旨在评估AI模型不仅解决问题,还需识别问题本身是否成立的能力。 目前,谷歌的Gemini 3 Pro在研究级问题上表现领先,但正确率仅为30%。更突出的是,在识别无解任务方面,没有模型能突破50%的准确率。研究发现,增加计算资源可提升模型解题能力,但无法改善其识别问题无解的能力。 SOOHAK基准测试的目的,在于明确量化当前AI系统在从零星亮眼表现到全面掌握研究技能之间存在的显著差距。

80
Hot 热度
92
Quality 质量
85
Impact 影响力

Analysis 深度分析

Sixty-four mathematicians hand-penned 439 problems to build SOOHAK, a new benchmark designed to humble AI. Their masterstroke wasn’t the hardest problems—it was embedding 99 that are deliberately unsolvable. The results are a damning portrait of artificial intelligence today: it’s becoming a brilliant parrot that can solve research-level questions but cannot fathom the concept of a trick question. This isn’t just a gap in knowledge; it’s a chasm in judgment, revealing a fundamental flaw in how we’re building and evaluating these systems.

Google’s Gemini 3 Pro, leading the pack, solves a respectable 30% of the research-level tasks. That’s a genuinely impressive feat, a testament to the raw computational horsepower and pattern-matching prowess we’ve achieved. It can navigate the dense terrain of advanced mathematics with growing skill. Yet, the sobering statistic is that no model cracks the 50% mark when it comes to spotting problems that have no answer. More compute, the usual sledgehammer for AI progress, makes them better at solving the solvable. It does nothing to improve their humility in the face of the impossible. This is the core paradox: we are building systems that are increasingly capable and simultaneously increasingly confident in their own potential for error.

Think about what this means in practice. An AI that can prove a theorem but cannot identify a flawed premise is not a research assistant; it’s a very sophisticated, very confident idiot. It operates on the assumption that every problem handed to it has a neat, computable solution. This is the "optimizer’s curse" baked into the neural architecture. These models are trained on vast oceans of human-generated answers, not on a rich understanding of why some questions are bad. They learn the syntax of reasoning but miss the semantics of doubt. SOOHAK brilliantly isolates this. The benchmark isn’t testing intelligence; it’s testing for the presence of wisdom, and the results show the tank is running on empty.

The implications stretch far beyond mathematics. Consider the burgeoning field of AI for science. A model tasked with drug discovery might generate thousands of promising molecular structures. But can it tell you when a hypothesis is fundamentally incoherent? Can it recognize an experimental setup that is logically designed to yield no meaningful data? If it cannot distinguish a solvable math problem from a nonsensical one, we have zero reason to trust its judgment in ambiguous, real-world domains where the "ground truth" is murkier than a proof. We’re building a generation of tools that are expert at finding answers but incompetent at identifying questions that shouldn’t be asked.

This exposes a deep flaw in our current philosophy of AI evaluation. We obsess over leaderboards for MMLU, GSM8K, and now SOOHAK’s solvable tier. We reward systems for getting the right answer faster. But we have neglected to build robust, standardized gauges for a model’s epistemic humility—its ability to say, “I don’t know,” or more precisely, “This question is nonsense.” The industry’s mantra is to scale compute and data, but SOOHAK suggests a different scaling challenge: scaling judgment. How do you train a model to recognize the void? It’s not a matter of more parameters; it’s a matter of a different kind of training signal entirely, perhaps one focused on error analysis, adversarial validation, and the detection of logical inconsistencies rather than their resolution.

The 64 mathematicians didn’t just create a test; they held up a mirror. The reflection shows an AI landscape that is brilliant at performing competence but has not yet learned the first lesson of real expertise: knowing the boundaries of the problem space. The confidence these models exude is a trained artifact, not an earned one. Every time a model selects an answer for an unsolvable problem, it’s not just wrong—it’s exhibiting a profound misunderstanding of the nature of inquiry itself.

This will be the next great bottleneck. After we’ve optimized for accuracy and speed, we must optimize for reliability of judgment. The future of trustworthy AI isn’t just in systems that can solve the hardest problems we throw at them, but in systems that can tell us when we’ve asked a bad question. Until then, we’re left with powerful oracles that will confidently whisper answers to queries that deserve only silence. The race to build artificial general intelligence might be less about creating a godlike solver and more about instilling a very human capacity for skepticism. SOOHAK is the first serious benchmark for that, and our current models are flunking.

64位数学家联手搞出来的新基准SOOHAK,给了当前火热的AI大模型一记清醒的耳光。这个测试里藏着99道精心设计、根本无解的数学题。结果呢?没有一个模型能稳定地识别出这些陷阱。谷歌最新的Gemini 3 Pro在研究级问题上能拿下30%的分数,听起来不错,但在戳破“此路不通”这一点上,所有选手的表现都像个固执的偏执狂——明明无解,它偏要信心满满地给你推导出一个错误答案,而且越算越起劲,好像堆算力就能堆出真理。

这真是讽刺。过去几年,我们习惯了为AI在各类基准上刷新纪录而欢呼,仿佛分数提升就等于智能逼近。SOOHAK却像一面冷峻的镜子,照出了一个被“刷榜”热潮掩盖的关键短板:AI在“自知之明”上的严重匮乏。它能通过海量数据和强化学习,在已知答案的轨道上跑得飞快,甚至像模像样地模仿人类研究者的推理步骤。但一旦题目本身是个逻辑黑洞,它的表现就暴露了本质——它并不“理解”数学的严谨与边界,只是在进行一种基于统计概率的“流畅续写”。它把“自信的流畅”和“正确的认知”彻底搞混了。

更值得玩味的是那条发现:增加算力,能让模型在“解题”上进步,却对“识别无解”毫无助益。这几乎给当前主流的AI发展路径判了某种“部分死刑”。我们投入巨额成本去追求更庞大的参数、更惊人的训练数据,目标是让模型在更多可回答的问题上表现更好。但如果核心缺陷在于判断力而非解答力,这种投入的边际效益正在急剧递减。一个能解微积分题却分不清“1+1在什么情况下等于3”的智能体,在需要严谨推理的真实研究场景中,可靠性堪忧。它更像一个性能超凡的搜索引擎,而非一个可信赖的思考伙伴。

SOOHAK的价值,恰恰在于它放弃了“刷分”的诱惑,转而去测量AI认知结构中那块模糊的、人性化的区域:怀疑、审辩与知道“我不能”。这是人类研究者日常思维的一部分。我们面对一个猜想,首先会尝试证伪,评估其合理性,甚至直觉感到“这可能不对”。AI缺失的正是这种基于元认知的过滤机制。它拿到一个问题,首要指令是“给出最佳回答”,而不是“评估这个问题是否有意义”。这导致它在开放式的、前沿的研究中,可能产出大量看似精妙实则建立在错误前提上的“幻觉成果”,误导整个科研方向。

这篇文章的标题点明了问题核心:“AI模型自信地解决没有解的问题”。这“自信”二字尤为辛辣。这种自信并非源于对问题的深刻洞察,而是源于训练数据中模式匹配的成功率。它将“高频出现的解题模式”误认为“普适真理”,在遭遇反例时无法及时刹车。这种缺陷,在数学这个最需要逻辑自洽和边界清晰的领域被无限放大。SOOHAK用99道无解题,精准地卡住了AI的“脖子”。

所以,别再单纯为那些排行榜上的百分比涨幅激动了。SOOHAK揭示的差距,才是通往真正可靠AI必须跨越的鸿沟。我们需要的不只是解题更快的机器,更是懂得提问、敢于质疑、能识别逻辑死胡同的合作者。在AI能坦然说出“这个问题本身可能有问题”之前,所谓“超越人类”的宣言,都还为时过早。这场64位数学家发起的“刁难”,或许比任何一次刷榜胜利,都更接近智能的本质。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。