Back to Blog
News
7 min read

Mathematicians Demand OpenAI Prove It Didn't Steal Their Work

Matt Weitzman
Senior SEO Strategist & Co-Founder
Mathematicians Demand OpenAI Prove It Didn't Steal Their Work

A second mathematician has publicly accused OpenAI of "dishonest" behavior and a troubling lack of transparency around its AI training data. According to Mathematicians want proof OpenAI didn't use their work, mathematician Andreas Thom raised the allegations just days after a separate dispute erupted over whether OpenAI's models had benefited from unpublished academic work — adding more pressure to a company that has been riding high on a string of high-profile mathematical breakthroughs.

Thom published his concerns in a series of posts on Mastodon, arguing that conversations he and his colleagues had with ChatGPT before OpenAI's public announcements may have quietly contributed to the company's results. One of the 10 results OpenAI announced last month involved Thom's area of expertise — so-called non-sofic groups — and OpenAI acknowledged the result built heavily on prior work by Thom and fellow mathematician Gábor Kun.

OpenAI did not immediately respond to The Verge's request for comment.

The Details: What Thom Is Actually Claiming

Thom said he began reflecting on his own situation after Tristan Buckmaster, a mathematics professor at New York University, publicly questioned whether OpenAI's AI models had benefited from his use of OpenAI's Codex tool. The parallel was hard to ignore.

What struck Thom most was, in his words, "OpenAI's detailed command of our techniques" — approaches he said were neither the most obvious nor the most promising routes to a solution at the time. That level of specific familiarity raised questions he felt he had to ask directly.

He emailed OpenAI researchers Sébastien Bubeck and Mark Sellke — also a statistician at Harvard — asking whether his ChatGPT interactions were "part of the training data or accessible to the reasoning process." The answer he received, he said, only addressed whether his conversations could be accessed directly. It did not address whether those conversations had entered OpenAI's training data pools. "No such qualification, explanation, or evidence was given," Thom wrote. "I take this as dishonesty to say the least."

Thom's core argument is straightforward and hard to dismiss: researchers don't have the tools to reverse-engineer OpenAI's training pipeline. "Only OpenAI has the relevant data for that." He placed the burden squarely on the company — if OpenAI wants to deny using researcher interactions as training material, it should disclose the relevant datasets and clarify its data usage terms.

OpenAI's response to Buckmaster's earlier complaint followed a similar pattern. In its blog post announcing a solution to the Navier-Stokes problem — a famous fluid dynamics challenge — the company flatly denied accessing specific user data. But it stopped short of ruling out indirect influence, writing: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models." Thom called that the same "obfuscatory distinction" it drew with him. His rebuttal was pointed: "De-identification may remove a name; it does not remove the intellectual content of a mathematical idea."

Why This Matters (And Not Just for Mathematicians)

Think about what's actually being alleged here. Researchers share sophisticated, unpublished thinking with an AI chatbot. That thinking — stripped of names, but not of ideas — potentially improves the very model that then races them to a discovery they were working toward. No consent. No credit. No disclosure. Thom called this scenario "ethically indefensible" if true.

The incident is already changing behavior. According to The Verge's reporting, numerous researchers said they worry this kind of dynamic will push mathematics into a more secretive state — where even a rumor of progress could trigger a race with a well-resourced tech giant chasing glory.

That chilling effect has implications far beyond academia. If professionals across any knowledge-intensive field stop sharing authentic, expert thinking with AI tools out of fear it will be used against them, the quality of training data degrades over time. The feedback loop that makes these models increasingly capable breaks down. OpenAI's reluctance to give clear answers may be legally cautious, but it's strategically shortsighted.

For anyone in SEO or digital marketing, this story is a signal worth watching. The question of what goes into AI training data — and how that shapes AI-generated outputs — is not abstract. It has direct bearing on how AI Overviews, Perplexity, and ChatGPT surface information and whom they attribute it to. If AI companies are quietly absorbing expert knowledge without disclosure, the entire premise of E-E-A-T as a differentiator gets murkier.

What to Do Now

You probably can't audit OpenAI's training pipeline. Nobody outside that building can. But you can make smarter decisions about how you and your clients interact with AI tools going forward. Here's where to focus:

  1. Audit what proprietary knowledge your team shares with AI chatbots. Unpublished strategies, unreleased client data, internal research — understand what you're sending into these systems and under what terms.
  2. Read the data usage policies for any AI tool you're using professionally. Most major platforms allow users to opt out of having conversations used for training. Check whether you've actually done that.
  3. Treat original expert content as your defensible asset. If AI models are absorbing and redistributing expert thinking, the answer isn't to stop producing it — it's to publish it publicly and attributably first, so the authorship is clear.
  4. Monitor how AI tools answer questions in your niche. If an AI Overview or ChatGPT response sounds suspiciously close to your unpublished methodology, that's worth documenting. It's early days for enforcement, but documentation matters.
  5. Watch how this dispute resolves. If OpenAI is forced to disclose more about its training data sources — through regulatory pressure or public accountability — the implications for content creators and marketers will be significant.

Background: A Pattern of Opaque Data Practices

This isn't OpenAI's first brush with training data controversy, and the math community isn't its first critic. What's different here is the specificity and the stakes. Thom isn't making a vague copyright claim — he's pointing to a specific result in a specific field that directly overlaps with work he was doing, and he's naming the researchers he contacted. That precision makes the allegations harder to bat away.

The non-sofic groups result was one of 10 results OpenAI announced, according to The Verge. After the announcement, the company was widely criticized in mathematical circles for failing to acknowledge recent contributions from Thom and Kun. OpenAI quietly amended its writeup after the criticism — which is a notable admission in itself, even if the company didn't frame it that way.

OpenAI's Navier-Stokes announcement — the solution to one of mathematics' legendary Millennium Prize problems — was described by The Verge as an extraordinary achievement that few would deny if verified. But the unusual circumstances the company cited for pursuing it added a layer of unease: OpenAI said it heard rumors online that other researchers had made major progress and decided to try as well. That framing alone has left a sour taste in many researchers' mouths.

I've watched AI transparency questions percolate through search marketing conferences for a couple of years now. What's shifting is that the people asking these questions are no longer just ethicists and journalists — they're domain experts who can make specific, technical accusations that are hard to hand-wave away. That changes the conversation significantly.

If you're tracking how AI companies use content in their models — and how that affects your visibility in AI-powered search results — tools like AI visibility tracking are worth keeping in your stack so you can spot shifts as they happen.

This dispute is unlikely to end with Thom's Mastodon posts. Expect more researchers to come forward, expect more regulatory attention, and expect OpenAI to face harder questions about what exactly goes into making its models this good at things they weren't good at a year ago. The math community is organized, vocal, and not inclined to let this slide.

Frequently Asked Questions

Matt Weitzman

About

Senior SEO Strategist & Co-Founder

Matt has over 15 years of experience in technical SEO and digital marketing. He specializes in algorithmic recovery, enterprise architecture, and leveraging AI for content scaling. He is a frequent speaker at search marketing conferences.

More articles by Matt Weitzman