AI News AI资讯 4h ago Updated 1h ago 更新于 1小时前 48

Mathematicians want proof OpenAI didn't use their work 数学家要求证明OpenAI未使用其研究成果

Mathematician Andreas Thom accuses OpenAI of using unpublished research from his conversations with ChatGPT to achieve breakthroughs in non-sofic groups, calling the company's denials "dishonest" OpenAI acknowledges its result built on Thom and Gábor Kun's prior work but initially failed to properly credit them, later amending its writeup Thom argues OpenAI's distinction between "direct access" and "de-identified training data" is misleading, as intellectual content survives de-identification Th 数学家Andreas Thom指控OpenAI可能将其与ChatGPT的对话内容纳入训练数据,用于生成非sofic群相关数学成果 OpenAI承认成果建立在Thom与Kun先前工作基础上,但最初未充分致谢,后私下修改声明 OpenAI无法排除"去标识化用户数据"间接提升模型的可能性,被指回避核心质疑 数学界担忧AI竞赛文化将迫使研究者转向更封闭的研究模式

72
Hot 热度
65
Quality 质量
70
Impact 影响力

Analysis 深度分析

TL;DR

  • Mathematician Andreas Thom accuses OpenAI of using unpublished research from his conversations with ChatGPT to achieve breakthroughs in non-sofic groups, calling the company's denials "dishonest"
  • OpenAI acknowledges its result built on Thom and Gábor Kun's prior work but initially failed to properly credit them, later amending its writeup
  • Thom argues OpenAI's distinction between "direct access" and "de-identified training data" is misleading, as intellectual content survives de-identification
  • The controversy echoes prior disputes with mathematician Tristan Buckmaster over OpenAI's Navier-Stokes solution, raising systemic concerns about consent and credit
  • Researchers fear the episode will drive the mathematics community toward secrecy, as sharing ideas online risks triggering races with well-resourced AI labs

Why It Matters

This case represents a growing flashpoint between AI companies and academic researchers over data provenance, consent, and attribution in the age of large-scale AI systems. It challenges the industry's standard practices around training data collection and raises urgent questions about whether AI companies should be required to disclose whether user interactions have entered their training pipelines.

Technical Details

  • The controversy centers on OpenAI's result involving non-sofic groups, an area of expertise of mathematician Andreas Thom and Gábor Kun; OpenAI admitted the result "built heavily" on their prior unpublished work
  • Thom questioned whether his prior ChatGPT interactions influenced the model's reasoning, specifically noting OpenAI's "detailed command" of techniques that were neither the most obvious nor most promising approaches at the time
  • OpenAI's response distinguished between direct access to user data (denied) and indirect influence through de-identified training data (not ruled out), a distinction Thom called "materially misleading"
  • The incident follows a similar dispute with Tristan Buckmaster regarding OpenAI's Navier-Stokes Millennium Prize claim, where the company made nearly identical statements about not accessing specific user data
  • Thom and colleagues are not equipped to reverse-engineer OpenAI's training pipeline; only the company possesses the data necessary to verify or refute the claims

Industry Insight

  • AI companies must develop transparent, verifiable frameworks for handling user-generated content in training data, as vague denials erode trust with academic and research communities
  • The "de-identified data" loophole OpenAI invokes could expose the industry to widespread legal and ethical challenges if researchers' unpublished ideas are absorbed into models without consent or attribution
  • Organizations racing to solve high-profile problems using AI should establish clear ethical guidelines around data provenance, credit attribution, and researcher consent before deploying models on cutting-edge research problems

TL;DR

  • 数学家Andreas Thom指控OpenAI可能将其与ChatGPT的对话内容纳入训练数据,用于生成非sofic群相关数学成果
  • OpenAI承认成果建立在Thom与Kun先前工作基础上,但最初未充分致谢,后私下修改声明
  • OpenAI无法排除"去标识化用户数据"间接提升模型的可能性,被指回避核心质疑
  • 数学界担忧AI竞赛文化将迫使研究者转向更封闭的研究模式

为什么值得看

本文揭示了AI系统训练数据透明度与学术伦理交叉的核心争议,直接影响AI研发者如何界定"合理使用"边界。OpenAI在重大科学突破中的数据来源模糊性,可能重塑学术界与AI企业的合作信任基础。

技术解析

  • 争议聚焦于非sofic群(infinite mathematical structures不可被有限结构逼近的领域)的AI辅助证明,OpenAI模型展现出对特定数学技巧的"详细掌握"
  • Thom通过Mastodon公开质疑,指出其2023年与ChatGPT的技术讨论可能早于OpenAI公开成果,但OpenAI仅回应"未直接访问对话内容",未证明数据是否进入训练池
  • OpenAI在Navier-Stokes问题声明中采用"去标识化数据可能间接改进模型"的模糊表述,Thom批评该说法"在材料上具有误导性"
  • 数学界缺乏反向工程OpenAI训练管道的能力,要求企业主动披露数据集构成与使用条款

行业启示

  • AI企业需在重大科学突破中建立可验证的数据溯源机制,避免"去标识化"成为规避伦理审查的漏洞
  • 学术界应制定AI辅助研究的署名与数据使用规范,防止"AI竞赛"导致知识生产的不平等
  • 企业透明度缺失可能引发监管反弹,建议主动公开训练数据采样策略与用户数据使用边界

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Closed Source 闭源 LLM 大模型 Dataset 数据集 Ethics 伦理 Research 科学研究