FULL STORY

OpenAI's Data Use Under Fire: From User Complaint to Attribution Fight

OpenAI's admission that anonymized product data may feed training sparked a researcher backlash, escalating into a broader fight over idea aggregation and academic attribution around its Navier-Stokes claims.

2026-09-09 ~ 2026-09-10 · 2 episodes · 16 posts

Episode 1 · Researchers Suspect OpenAI Secretly Trains on Their Inputs (2026-09-09, 13 posts)

After OpenAI admitted it "cannot rule out that de-identified product usage data is used for model training," a debate over training data ethics and academic credit quickly spread through the research community. The trigger was user irldanB's complaint that OpenAI could learn from over a year of collaboration with ChatGPT—code records he considered on par with Millennium Prize/Fields Medal level—and, amid rumors he was close to finishing, beat him to the result; he noted Anthropic had a similar case. Israeli NLP researcher Yoav Goldberg pointed out the key distinction: "I accept that a model trained on my data is released publicly" and "I accept an internal model trained on my latest work solving the problem first" are two entirely different things; Boaz Barak added that users can opt out of (de-identified) data being used for training.

Confirmed

  • OpenAI stated it cannot guarantee user conversations stay out of training data, nor rule out that de-identified product usage data helps improve models.
  • One user noted OpenAI began training an updated model on August 28, but discussants considered this unlikely to affect the research conclusions in question, with details involving model variants, fine-tuning, and internal use of models at different training stages.

Controversy and skepticism

  • Safety researcher davidmanheim demanded OpenAI tell the truth: either it has no tools to track what data it trained on (unacceptable), or it has the tools but refuses to disclose (worse).
  • Developer ctjlewis argued vendors should never promise conversations won't be used for training in the first place, and OpenAI's "excessive honesty" precisely exposes that it can't guarantee it; he asserted that anything anyone has said to GPT or published in the past few months has already gone into the last training round.
  • Researcher gleech commented that OpenAI's real situation may be "the data was used automatically, not intentionally, but it's impossible to prove it didn't affect the results—because deep learning is not a science," which he believes is most likely true but extremely damaging for OpenAI.

Why it matters

  • The dispute touches on the fundamental trust issue between AI vendors and users/researchers: the untraceability of training data makes "whether someone's work was used to scoop them" nearly impossible to falsify, posing a substantive threat to academic research that relies on open public collaboration.

Episode 2 · Navier-Stokes dispute raises questions of attribution for aggregated ideas (2026-09-09, 3 posts)

The Navier-Stokes dispute has sparked discussion about AI labs aggregating users' half-formed ideas during training, with researchers advising data-use opt-outs and highlighting attribution as the core unresolved problem.