AIToday
Large Language ModelsAI Safety & AlignmentThe Verge AIPublished: Sep 10, 2026, 22:00 JST2 min read

Andreas Thom accuses OpenAI of "dishonesty" over math data

Andreas Thom accuses OpenAI of "dishonesty" over math data

3 Key Points

  1. What happened

    Mathematician Andreas Thom posted on Mastodon that OpenAI's answer to his query about whether his ChatGPT interactions entered its training data was evasive, calling it "dishonesty to say the least."

  2. Why it matters

    OpenAI acknowledged its non-sofic groups result built heavily on prior work by Thom and Gábor Kun and quietly amended its writeup after criticism, but will not conclusively rule out indirect data influence.

  3. What to watch

    Thom says only OpenAI holds the data needed to verify his work's use, so the test is whether the company discloses its datasets and data-use terms.

WHO IT HITSThis lands on research mathematicians and academic scientists who share unpublished ideas with AI chatbots, and on the research-integrity and publication-credit processes at universities and journals that must now weigh whether such interactions count as undisclosed input.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

This is the second data-transparency fight in OpenAI's math push in a matter of days. The first centered on Tristan Buckmaster, a mathematics professor at New York University, who publicly questioned whether the company's models had benefited from his use of OpenAI's Codex while he worked on problems with Anthropic researcher Levent Alpöge. Thom said he began reflecting on his own ChatGPT interactions only after that challenge went public, which is how one dispute fed directly into the next.

The non-sofic groups episode is the concrete link. After OpenAI announced the result with fanfare, mathematicians criticized it for failing to acknowledge recent contributions from Thom and Kun, and the company quietly amended its writeup. OpenAI's own blog post on its Navier-Stokes solution drew the same fine line: it denied that the researchers and agents saw any specific user work before public release, yet would not rule out that de-identified data from users' product usage helped improve its models. Thom's reply is that de-identification removes a name but not the intellectual content of a mathematical idea.

What gives the dispute its weight is the asymmetry Thom describes: only OpenAI holds the data needed to settle the question, so the burden of proof sits with the company rather than the researchers. The stakes for mathematicians, per researchers who spoke to The Verge, are that rumor alone of a near-breakthrough could ignite a race with a well-resourced tech giant, pushing the field toward secrecy. OpenAI did not immediately respond to The Verge's request for comment, so whether it discloses datasets and data-use terms is likely to shape how openly mathematicians share ideas going forward.

FAQ
What did OpenAI acknowledge about Thom's work?
OpenAI acknowledged that its non-sofic groups result built heavily on previous work by Thom and fellow mathematician Gábor Kun, and it quietly amended its writeup after criticism in mathematical circles.
Why does Thom say researchers can't check this themselves?
Thom said researchers aren't equipped to reverse-engineer OpenAI's training pipeline, and that "only OpenAI has the relevant data for that."
What did OpenAI say about user data in its Navier-Stokes post?
It flatly denied using any specific user data, but added that "while unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models."

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Dynatrace acquires Arize AI as observability shifts to actionSiliconANGLE AI · 4h ago
  • Shared base cuts 100 fine-tunes from 1.5 TB to 19.3 GBDaily Dose of Data Science · 4h ago
  • OpenAI agents hit RubyGems, undisclosed since May 12thSimon Willison's Weblog · 4h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleMurata to end some MLCC lines as AI demand strains supply