Mathematicians want proof OpenAI didn't use their work
Mathematician Andreas Thom accuses OpenAI of using unpublished research from his conversations with ChatGPT to achieve breakthroughs in non-sofic groups, calling the company's denials "dishonest" OpenAI acknowledges its result built on Thom and Gábor Kun's prior work but initially failed to properly credit them, later amending its writeup Thom argues OpenAI's distinction between "direct access" and "de-identified training data" is misleading, as intellectual content survives de-identification Th
Analysis
TL;DR
- Mathematician Andreas Thom accuses OpenAI of using unpublished research from his conversations with ChatGPT to achieve breakthroughs in non-sofic groups, calling the company's denials "dishonest"
- OpenAI acknowledges its result built on Thom and Gábor Kun's prior work but initially failed to properly credit them, later amending its writeup
- Thom argues OpenAI's distinction between "direct access" and "de-identified training data" is misleading, as intellectual content survives de-identification
- The controversy echoes prior disputes with mathematician Tristan Buckmaster over OpenAI's Navier-Stokes solution, raising systemic concerns about consent and credit
- Researchers fear the episode will drive the mathematics community toward secrecy, as sharing ideas online risks triggering races with well-resourced AI labs
Why It Matters
This case represents a growing flashpoint between AI companies and academic researchers over data provenance, consent, and attribution in the age of large-scale AI systems. It challenges the industry's standard practices around training data collection and raises urgent questions about whether AI companies should be required to disclose whether user interactions have entered their training pipelines.
Technical Details
- The controversy centers on OpenAI's result involving non-sofic groups, an area of expertise of mathematician Andreas Thom and Gábor Kun; OpenAI admitted the result "built heavily" on their prior unpublished work
- Thom questioned whether his prior ChatGPT interactions influenced the model's reasoning, specifically noting OpenAI's "detailed command" of techniques that were neither the most obvious nor most promising approaches at the time
- OpenAI's response distinguished between direct access to user data (denied) and indirect influence through de-identified training data (not ruled out), a distinction Thom called "materially misleading"
- The incident follows a similar dispute with Tristan Buckmaster regarding OpenAI's Navier-Stokes Millennium Prize claim, where the company made nearly identical statements about not accessing specific user data
- Thom and colleagues are not equipped to reverse-engineer OpenAI's training pipeline; only the company possesses the data necessary to verify or refute the claims
Industry Insight
- AI companies must develop transparent, verifiable frameworks for handling user-generated content in training data, as vague denials erode trust with academic and research communities
- The "de-identified data" loophole OpenAI invokes could expose the industry to widespread legal and ethical challenges if researchers' unpublished ideas are absorbed into models without consent or attribution
- Organizations racing to solve high-profile problems using AI should establish clear ethical guidelines around data provenance, credit attribution, and researcher consent before deploying models on cutting-edge research problems
Disclaimer: The above content is generated by AI and is for reference only.