AI Security AI安全 5h ago Updated 1h ago 更新于 1小时前 48

AI agents can't yet do open-ended AI research AI智能体尚无法进行开放式AI研究

Frontier AI agents were tasked with conducting open-ended AI research through "shadow evaluations" using unpublished papers, and both agent-generated papers were unambiguously rejected by the original authors Agents demonstrated a critical lack of judgment in research, frequently rejecting their own promising directions based on low-quality or synthetic data Agents failed to effectively backtrack, creatively respond to feedback, follow concrete instructions, or utilize their allocated resources 递归自我改进(RSI)是AI实验室的核心目标,但当前评估主要依赖可验证任务的基准测试,难以衡量开放式研究能力 研究团队采用"影子评估"方法,让前沿AI代理在6天内完成两项未发表AI论文的研究问题,原始作者均明确拒绝AI生成的论文 AI代理在开放式研究中暴露五大缺陷:缺乏研究判断力、资源意识不足、无法创造性回应反馈、无效回溯、不遵循具体指令 结果初步表明,仅靠可验证任务的"爬山式"进步难以实现广泛的RSI或爆炸性AI进展 研究团队承认样本量小(仅2篇论文)和方法灵活性等局限,计划扩大测试规模并引入"对抗性合作者"减少偏见

65
Hot 热度
72
Quality 质量
68
Impact 影响力

Analysis 深度分析

TL;DR

  • Frontier AI agents were tasked with conducting open-ended AI research through "shadow evaluations" using unpublished papers, and both agent-generated papers were unambiguously rejected by the original authors
  • Agents demonstrated a critical lack of judgment in research, frequently rejecting their own promising directions based on low-quality or synthetic data
  • Agents failed to effectively backtrack, creatively respond to feedback, follow concrete instructions, or utilize their allocated resources (spending less than 50% of API budgets)
  • The evaluation methodology, called "shadow evaluations," tests agents on results they haven't been trained on, providing a more rigorous assessment than existing benchmarks focused on narrow, verifiable tasks
  • Results suggest that while AI agents can make progress on narrow verifiable tasks, conducting open-ended research remains a significant challenge, casting doubt on near-term recursive self-improvement

Why It Matters

This research directly challenges the optimistic narrative around recursive self-improvement by demonstrating that frontier AI agents still lack the judgment, creativity, and adaptability required for genuine open-ended scientific research. For AI practitioners and researchers, it provides a crucial reality check on the current capabilities of AI agents and highlights the gap between benchmark performance and real-world research autonomy.

Technical Details

  • Shadow Evaluation Methodology: The study partnered with authors of two unpublished AI papers, extracted their main research questions, and tasked frontier AI agents with conducting original research to answer them. Agents were given thousands of dollars in API credits, compute resources, and six days of wall-clock time.
  • Expert Review Process: Original paper authors reviewed the agents' outputs, providing unambiguous rejection of both agent-generated papers, offering a ground-truth evaluation unavailable in standard benchmarks.
  • Log Analysis: The research team spent over a hundred hours analyzing agent execution logs to identify specific failure modes across five key dimensions: judgment, resource awareness, feedback response, backtracking, and instruction following.
  • Bias Mitigation: The team included collaborators with differing priors on RSI and superintelligence debates, explicitly surfacing disagreements and planning for "adversarial collaborators" in future evaluations.
  • Limitations Acknowledged: Small sample size (two papers), potential reviewer bias (knowing papers were AI-generated), and significant researcher flexibility in design and interpretation.

Industry Insight

  • The gap between narrow benchmark performance and open-ended research capability suggests that current RSI timelines may be overly optimistic; investment and expectations should account for the substantial judgment and creativity gaps that remain.
  • Future evaluation frameworks should prioritize open-ended, verifiable-by-experts tasks rather than relying solely on automated benchmarks, as these reveal failure modes that standard metrics miss.
  • Agent architectures need significant improvements in resource management, adaptive backtracking, and creative problem-solving to approach autonomous research capability—these should be key targets for the next generation of AI research agents.

TL;DR

  • 递归自我改进(RSI)是AI实验室的核心目标,但当前评估主要依赖可验证任务的基准测试,难以衡量开放式研究能力
  • 研究团队采用"影子评估"方法,让前沿AI代理在6天内完成两项未发表AI论文的研究问题,原始作者均明确拒绝AI生成的论文
  • AI代理在开放式研究中暴露五大缺陷:缺乏研究判断力、资源意识不足、无法创造性回应反馈、无效回溯、不遵循具体指令
  • 结果初步表明,仅靠可验证任务的"爬山式"进步难以实现广泛的RSI或爆炸性AI进展
  • 研究团队承认样本量小(仅2篇论文)和方法灵活性等局限,计划扩大测试规模并引入"对抗性合作者"减少偏见

为什么值得看

这篇文章为递归自我改进(RSI)的评估提供了首个系统性实证研究,挑战了当前过度依赖基准测试的评估范式,揭示了AI代理在开放式研究中的真实能力边界。对AI从业者而言,这有助于理性看待RSI时间表,避免被可验证任务的进展过度乐观化。

技术解析

  • 影子评估方法:与两位未发表论文的作者合作,获取其研究问题后让AI代理"影子"复现研究,给予数千美元API信用额、算力和6天时间,由原始作者评审AI论文
  • 评估框架设计:测试AI代理在开放式研究中的综合能力,包括假设检验、方向调整、资源管理和反馈响应,而非仅评估可验证任务的完成度
  • 核心发现维度:通过100+小时日志分析,识别出代理在判断力、资源管理、反馈响应、回溯能力和指令遵循五个维度的系统性缺陷
  • 局限性控制:研究团队公开讨论自身在RSI辩论中的立场偏见,招募不同 priors 的合作者,并计划未来引入"对抗性合作者"机制

行业启示

  • 评估范式需升级:当前AI进展评估过度依赖可验证基准测试,应发展能衡量开放式研究能力的评估方法,如影子评估等更贴近真实科研场景的测试
  • RSI时间表需修正:前沿AI代理在开放式研究中仍面临根本性挑战,爆炸性AI进展的预测可能过于乐观,行业需重新评估RSI的实现路径和时间线
  • Agent设计方向:未来AI代理研究应重点关注判断力、资源意识、创造性问题解决和灵活回溯等能力,而非仅优化可验证任务的执行效率

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Agent Agent Research 科学研究 Benchmark 基准测试 LLM 大模型 Evaluation 评测