AI News AI资讯 1d ago Updated 23h ago 更新于 23小时前 49

OpenAI's millennium proof dispute raises the question of whether researchers can trust AI labs OpenAI千年难题证明争议引发研究者能否信任AI实验室的疑问

Researcher Tristan Buckmaster accuses OpenAI of using his and Levent Alpöge's drafts (uploaded to Codex) to train its model on the Navier-Stokes Millennium Problem, alleging plagiarism, pressure tactics, and the exclusion of Alpöge (an Anthropic employee) as a co-author OpenAI CEO Sam Altman and researcher Sébastien Bubeck deny the allegations, claiming the team "acted with integrity" and that the effort was triggered by rumors that Anthropic's models had solved the problem OpenAI acknowledges i OpenAI被数学家Tristan Buckmaster指控在Navier-Stokes方程(千禧年难题)证明上存在学术不端,包括使用其草稿训练模型、施压研究者、试图移除Anthropic员工Alpöge作为合著者 OpenAI CEO Sam Altman和研究员Sébastien Bubeck否认抄袭指控,但承认在听到"Anthropic模型可能解决千禧年难题"的传言后,调动大量资源攻关同一问题 数学家Terence Tao警告:仅凭传言就能触发科技公司调动海量AI资源"压平"原创研究,可能逆转开放科学传统 研究者向OpenAI系统输入数据存在被自身成果超越的风险,现有"排除训练"选项保护

72
Hot 热度
68
Quality 质量
70
Impact 影响力

Analysis 深度分析

TL;DR

  • Researcher Tristan Buckmaster accuses OpenAI of using his and Levent Alpöge's drafts (uploaded to Codex) to train its model on the Navier-Stokes Millennium Problem, alleging plagiarism, pressure tactics, and the exclusion of Alpöge (an Anthropic employee) as a co-author
  • OpenAI CEO Sam Altman and researcher Sébastien Bubeck deny the allegations, claiming the team "acted with integrity" and that the effort was triggered by rumors that Anthropic's models had solved the problem
  • OpenAI acknowledges it "cannot rule out" that de-identified data from user inputs may have helped improve its models, despite employees downplaying the likelihood of training on the researchers' solutions
  • Mathematician Terence Tao warns that the incident could reverse centuries of open science traditions, as researchers may stop sharing promising directions for fear that tech companies will mobilize massive resources based on rumors and overtake their work

Why It Matters

This dispute strikes at the heart of trust between the academic research community and AI labs, raising urgent questions about data provenance, consent, and the ethics of training on user-submitted research. It also highlights a structural power imbalance: a single rumor can trigger a well-funded lab to deploy massive compute resources and potentially publish first, undermining the incentive structure that has long sustained open scientific collaboration.

Technical Details

  • The dispute centers on the Navier-Stokes equations, one of the seven Clay Mathematics Institute Millennium Problems carrying a $1 million prize, and involves claims that an AI system produced a significant mathematical result on this problem
  • Buckmaster and Alpöge allege they had entered a similar approach into OpenAI's Codex system, and that this input likely ended up in the training data used to develop OpenAI's model
  • OpenAI employee Boaz Barak disputed the necessity of external input, claiming the model independently proved a stronger result and did not need "hints" from the researchers
  • OpenAI's official blog post concedes uncertainty, stating that while unlikely, it cannot rule out that de-identified data derived from user product usage helped improve its models
  • The incident underscores the opacity of AI training pipelines: there is no public verification of whether the researchers' opt-out settings for data training were effective or even available

Industry Insight

  • AI labs must establish transparent, verifiable data governance policies if they wish to maintain credibility with the research community; vague disclaimers about "de-identified data" are insufficient to prevent erosion of trust
  • Researchers should treat any input into commercial AI systems as potentially contributing to competitive training data, and reconsider sharing preliminary results through such platforms until clearer legal and ethical safeguards exist
  • The incident signals a broader trend where AI labs operate as rapid-response research entities, capable of pivoting massive resources based on rumors—a dynamic that could fundamentally reshape academic publishing, priority claims, and the economics of scientific discovery

TL;DR

  • OpenAI被数学家Tristan Buckmaster指控在Navier-Stokes方程(千禧年难题)证明上存在学术不端,包括使用其草稿训练模型、施压研究者、试图移除Anthropic员工Alpöge作为合著者
  • OpenAI CEO Sam Altman和研究员Sébastien Bubeck否认抄袭指控,但承认在听到"Anthropic模型可能解决千禧年难题"的传言后,调动大量资源攻关同一问题
  • 数学家Terence Tao警告:仅凭传言就能触发科技公司调动海量AI资源"压平"原创研究,可能逆转开放科学传统
  • 研究者向OpenAI系统输入数据存在被自身成果超越的风险,现有"排除训练"选项保护有限

为什么值得看

本文揭示了AI实验室与学术界的信任危机,对AI从业者而言是理解大模型训练数据伦理边界的典型案例。对行业而言,它警示了科技公司凭借资源不对称可能颠覆传统科研竞争规则,影响整个AI研究生态的开放性。

技术解析

  • 争议核心:OpenAI是否将研究者上传至Codex的草稿纳入模型训练数据。OpenAI承认"无法排除去标识化数据可能帮助改进模型",但员工Boaz Barak认为模型本身能力足够,不需要外部"提示"
  • 千禧年难题背景:Navier-Stokes方程是Clay数学研究所设立的七个千禧年难题之一,悬赏100万美元。该问题研究流体动力学基本方程的正则性
  • 数据训练争议:研究者称已输入类似方法至OpenAI系统,怀疑输入成为训练数据。OpenAI员工指出若研究者未禁用训练数据收集选项,则实际训练可能性较低,但具体设置未公开
  • 作者归属分歧:Alpöge因任职Anthropic被排除在OpenAI论文作者之外,双方对此存在叙事差异——Altman称Alpöge不愿合作,Alpöge则表示愿意且作者身份不重要

行业启示

  • 开放科学面临结构性威胁:科技公司可仅凭"传言"就调动远超学术团队的资源进行攻关,可能使原创研究者失去优先发表权,颠覆数百年学术竞争规则
  • 数据输入风险需重新评估:研究者向AI系统输入未发表成果存在被"反向利用"的风险,现有opt-out机制保护力度不足,学术界需建立更严格的数据隔离规范
  • AI实验室需重建信任机制:OpenAI等公司在训练数据使用、研究者合作透明度方面缺乏清晰边界,行业亟需建立可验证的数据伦理框架以避免类似争议

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Research 科学研究 Ethics 伦理 Closed Source 闭源