deniz.in

Markets

Weather

Loading weather

· via The Verge

Second mathematician accuses OpenAI of dishonesty over training data origins

Mathematician Andreas Thom says OpenAI never demonstrated that his ChatGPT conversations were kept out of its training data, deepening a dispute over the provenance of its math breakthroughs.

Second mathematician accuses OpenAI of dishonesty over training data origins

A second mathematician has publicly accused OpenAI of unethical and dishonest behavior over where its training data comes from. According to The Verge, Andreas Thom wrote in a series of Mastodon posts that interactions he and his colleagues had with ChatGPT before OpenAI's recent string of mathematical announcements may have contributed to the company's results, and that OpenAI has declined to prove otherwise.

What Thom is alleging

One of the ten results OpenAI publicized last month sits in Thom's own specialty: non-sofic groups, which can loosely be described as infinite mathematical structures that cannot be approximated by finite ones. The Verge reports that OpenAI acknowledged its result rested heavily on earlier work by Thom and Gábor Kun, and that after the company was criticized in mathematical circles for failing to credit those contributions, it quietly amended its writeup.

Thom said he began scrutinizing his own dealings with the company after Tristan Buckmaster, a mathematics professor at New York University, publicly questioned whether OpenAI's models had gained an edge from his use of its Codex tool. Thom was also struck, he said, by OpenAI's detailed command of the techniques he and Kun had developed — routes to a solution that were neither the most obvious nor the most promising ones at the time.

An answer that avoided the question

Thom said he emailed OpenAI researchers Sébastien Bubeck and Mark Sellke, who is also a statistician at Harvard, asking whether his conversations with the chatbot were part of the training data or accessible to the reasoning process behind the result. The reply, according to Thom, addressed only whether his chats could be accessed directly, not whether they had entered the vast pools of data OpenAI uses to improve its models. "No such qualification, explanation, or evidence was given," he wrote, calling it dishonesty at minimum, and later describing Sellke's original answer as unjustifiably broad and materially misleading.

His central point is about asymmetry: researchers have no practical way to reverse-engineer OpenAI's training pipeline to check whether their material was absorbed. Only OpenAI holds the relevant data, Thom argues, so the obligation falls on the company to disclose the necessary datasets and clarify the settings and terms that govern how user data is handled.

The same pattern as the Millennium Prize dispute

The fine distinction Thom describes — no direct access, but no ruling out of indirect influence — mirrors how OpenAI defended its announced Navier-Stokes solution, a Millennium Prize problem concerning fluid motion, both publicly and in exchanges with Buckmaster, who had been working on the problem alongside Anthropic researcher Levent Alpöge in a personal capacity. In its announcement, OpenAI flatly denied that its researchers or agents had seen the pair's work before public release and that no specific user data was accessed. Yet it also conceded that while unlikely, it could not rule out that de-identified data derived from their product usage had helped improve its models.

Thom rejected that framing outright: de-identification may strip out a name, he argued, but it does not remove the intellectual content of a mathematical idea. He said it would be ethically indefensible if nonpublic research supplied by users ended up sharpening models that the company then used to beat those very users to publication, without consent, disclosure or credit. OpenAI did not immediately respond to The Verge's request for comment.

Why it matters

The controversy has landed at what should be a moment of celebration for OpenAI. Its claimed solution to a legendary Millennium Prize problem is, if verified, a remarkable achievement. But The Verge notes it was pursued under unusual circumstances: the company says it heard rumors online that other researchers had made major progress and decided to chase the problem itself.

That backdrop is feeding anxiety across mathematics. Several researchers told The Verge they fear that knowing a rumor of near-breakthrough could trigger a race with a well-resourced tech giant will push the field toward secrecy, undermining the open exchange that has long driven progress. The broader issue is accountability: if conversations with tools like ChatGPT can silently become training signal, then every researcher using these systems is potentially feeding a competitor built partly on their own unpublished ideas — and only the AI company itself can inspect the evidence of whether that happened.

  • #openai
  • #chatgpt
  • #training-data
  • #mathematics
  • #ai-ethics

Related posts