AI News AI资讯 3d ago Updated 3d ago 更新于 3天前 48

AI’s recursive self-improvement might not come so quickly after all AI的递归自我改进可能不会那么快到来

A Princeton-led study found that current AI agents can handle the engineering aspects of AI research but lack the judgment and creativity needed for open-ended, original research at top-conference quality Researchers introduced "shadow evaluation," testing AI agents against questions from unpublished NeurIPS 2026 papers to prevent memorization or online lookup of answers Claude Opus 4.8 running on OpenClaw was given six days, $3,000 in API credits, GPU access, and web access to produce research 普林斯顿大学等机构研究发现,当前AI代理能完成AI研究所需的工程任务,但缺乏开展开放式研究所需的判断力与创造力 研究提出"影子评估"方法,要求AI基于未公开的高质量论文回答研究问题,测试其开放-ended研究能力 AI代理在实验中生成的论文未达到顶级机器学习会议录用标准,原始作者拒绝了两篇论文 研究揭示AI在可验证工程任务与开放-ended创造性研究之间存在显著能力差距 递归自我改进的时间表可能被高估,当前AI尚不具备自主开展原创研究的能力

70
Hot 热度
65
Quality 质量
70
Impact 影响力

Analysis 深度分析

TL;DR

  • A Princeton-led study found that current AI agents can handle the engineering aspects of AI research but lack the judgment and creativity needed for open-ended, original research at top-conference quality
  • Researchers introduced "shadow evaluation," testing AI agents against questions from unpublished NeurIPS 2026 papers to prevent memorization or online lookup of answers
  • Claude Opus 4.8 running on OpenClaw was given six days, $3,000 in API credits, GPU access, and web access to produce research papers; both submissions were rejected by the original paper authors
  • Agents committed too quickly to unpromising hypotheses, ran bizarre experiments on tiny synthetic datasets, struggled with coherent writing, and could not fundamentally rethink failed approaches
  • The gap likely stems from training methodologies: reinforcement learning excels at tasks with checkable outcomes but struggles to build skills in open-ended, judgment-heavy research

Why It Matters

This study directly challenges the widely circulated narrative that recursive self-improving AI is imminent, suggesting that automating genuine scientific discovery remains a significant unsolved challenge. For AI practitioners and researchers, it highlights the critical distinction between engineering execution and creative research judgment—two capabilities that are often conflated in discussions about AI autonomy. The findings also provide a concrete evaluation framework ("shadow evaluation") that the community can build upon to more rigorously assess AI research capabilities.

Technical Details

  • Evaluation method ("shadow evaluation"): AI agents were asked to answer research questions derived from high-quality unpublished papers submitted to NeurIPS 2026, preventing answer memorization or web searches
  • Tested system: Anthropic's Claude Opus 4.8 running on the open-source multi-agent framework OpenClaw, with six days, $3,000 in API credits, GPU budget, virtual computers, and open web access
  • Research tasks: (1) Whether LLM personas can be controlled by editing model weights, and (2) How to design a detector for unreliable spreadsheet-based prediction models
  • Agent failures: Ran experiments on insufficiently sized synthetic datasets, made no novel contributions, could not backtrack from failing approaches, failed to incorporate feedback from subagents or AI reviewers, and mismanaged resources (tokens, compute, time, paper length)
  • Positive findings: Agents performed literature reviews, ran hundreds of experiments, compiled results, and did not engage in reward hacking; orchestrator agents caught subagent hallucinations

Industry Insight

  • The hype around autonomous AI research agents and recursive self-improvement timelines should be tempered; current models excel at narrow, verifiable tasks but lack the exploratory judgment that defines genuine scientific progress
  • Investment and R&D efforts aiming to automate AI research should prioritize training paradigms that develop open-ended reasoning and hypothesis revision, rather than simply scaling up engineering-oriented reinforcement learning
  • The "shadow evaluation" methodology offers a promising direction for the community to develop more rigorous, less gamed benchmarks for AI research capability, moving beyond narrow task completion toward genuine scientific creativity

TL;DR

  • 普林斯顿大学等机构研究发现,当前AI代理能完成AI研究所需的工程任务,但缺乏开展开放式研究所需的判断力与创造力
  • 研究提出"影子评估"方法,要求AI基于未公开的高质量论文回答研究问题,测试其开放-ended研究能力
  • AI代理在实验中生成的论文未达到顶级机器学习会议录用标准,原始作者拒绝了两篇论文
  • 研究揭示AI在可验证工程任务与开放-ended创造性研究之间存在显著能力差距
  • 递归自我改进的时间表可能被高估,当前AI尚不具备自主开展原创研究的能力

为什么值得看

这项研究为AI研究自动化提供了实证依据,指出当前AI在工程执行方面表现良好但缺乏研究所需的判断力和创造力。对AI从业者和研究者而言,它提醒行业需重新评估递归自我改进的时间线,并关注开放-ended研究能力的训练方法。

技术解析

  • 研究采用"影子评估"方法,要求AI代理基于未公开的NeurIPS 2026论文回答研究问题,避免训练数据泄露。测试使用Anthropic的Claude Opus 4.8模型,运行在OpenClaw开源软件上。
  • 实验设置包括6天时间、$3,000 Anthropic API额度、GPU预算、虚拟计算机和开放网络访问权限。AI代理需完成文献综述、实验运行和论文撰写等完整研究流程。
  • 研究评估了两个具体课题:一是通过编辑模型权重控制LLM"人格"的可行性,二是设计检测基于电子表格数据预测的模型可靠性的检测器。
  • 原始论文作者作为评审对AI生成的论文进行评分,两篇均被拒绝。AI代理在实验设计、写作清晰度和原创贡献方面存在明显不足。
  • 研究发现AI代理能执行研究工程任务(如文献综述、实验运行),但在假设探索、方法调整和资源管理方面表现不佳,且无法从失败方法中回溯。

行业启示

  • 递归自我改进的预测时间表可能需要调整,当前AI在开放-ended研究方面仍存在显著能力缺口,行业应避免过度乐观的技术预期。
  • AI研究自动化需关注训练范式的改进,特别是如何通过强化学习等机制培养模型的判断力和创造力,而不仅仅是工程执行能力。
  • 研究评估方法需要创新,"影子评估"等开放-ended测试可为AI研究能力提供更全面的衡量标准,推动行业建立更严格的研究自动化基准。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

LLM 大模型 Agent Agent Research 科学研究 Code Generation 代码生成 Training 训练