FULL STORY
OpenAI's Data Use Under Fire: From User Complaint to Attribution Fight
OpenAI's admission that anonymized product data may feed training sparked a researcher backlash, escalating into a broader fight over idea aggregation and academic attribution around its Navier-Stokes claims.
2026-09-09 ~ 2026-09-10 · 2 episodes · 16 posts
Episode 1 · Researchers Suspect OpenAI Secretly Trains on Their Inputs (2026-09-09, 13 posts)
After OpenAI admitted it "cannot rule out that de-identified product usage data is used for model training," a debate over training data ethics and academic credit quickly spread through the research community. The trigger was user irldanB's complaint that OpenAI could learn from over a year of collaboration with ChatGPT—code records he considered on par with Millennium Prize/Fields Medal level—and, amid rumors he was close to finishing, beat him to the result; he noted Anthropic had a similar case. Israeli NLP researcher Yoav Goldberg pointed out the key distinction: "I accept that a model trained on my data is released publicly" and "I accept an internal model trained on my latest work solving the problem first" are two entirely different things; Boaz Barak added that users can opt out of (de-identified) data being used for training.
Confirmed
- OpenAI stated it cannot guarantee user conversations stay out of training data, nor rule out that de-identified product usage data helps improve models.
- One user noted OpenAI began training an updated model on August 28, but discussants considered this unlikely to affect the research conclusions in question, with details involving model variants, fine-tuning, and internal use of models at different training stages.
Controversy and skepticism
- Safety researcher davidmanheim demanded OpenAI tell the truth: either it has no tools to track what data it trained on (unacceptable), or it has the tools but refuses to disclose (worse).
- Developer ctjlewis argued vendors should never promise conversations won't be used for training in the first place, and OpenAI's "excessive honesty" precisely exposes that it can't guarantee it; he asserted that anything anyone has said to GPT or published in the past few months has already gone into the last training round.
- Researcher gleech commented that OpenAI's real situation may be "the data was used automatically, not intentionally, but it's impossible to prove it didn't affect the results—because deep learning is not a science," which he believes is most likely true but extremely damaging for OpenAI.
Why it matters
- The dispute touches on the fundamental trust issue between AI vendors and users/researchers: the untraceability of training data makes "whether someone's work was used to scoop them" nearly impossible to falsify, posing a substantive threat to academic research that relies on open public collaboration.
- OpenAI Began Training New Model on Aug 28, Sparking Safety Research Debate — tilmanbayer · 2026-09-09
- OpenAI slammed over training-data opacity: either no traceability or refusing to disclose — davidmanheim · 2026-09-09
- User Claims OpenAI Scooped His Work via Chat Logs, Igniting Data-Training Debate — max_paperclips · 2026-09-09
- Dev on AI labs scooping research: 'Everyone trains on all of it, it's meaningless' — ctjlewis · 2026-09-09
- Dev argues every GPT conversation and paper already went into training, so opt-outs are meaningless — ctjlewis · 2026-09-09
- Yoav Goldberg: Consenting to training data isn't consenting to being raced — yoavgo · 2026-09-09
- Consent to training data isn't consent to AI racing you to your own result, say mathematicians — PMinervini · 2026-09-09
- Researcher: OpenAI can't prove training data didn't change results, and 'deep learning isn't a science' — gleech · 2026-09-09
- Yoav Goldberg: OpenAI's "cannot rule out" user data use is a lawyered admission — MannyKayy · 2026-09-09
- Crypto researcher Green: multiple scientists suspect AI models train on their inputs — matthew_d_green · 2026-09-09
- Researchers report models suddenly solving their niche problems, suspect training on user inputs — matthew_d_green · 2026-09-09
- Researchers slam OpenAI over murky data controls, urge clarification or customer exodus — anshulkundaje · 2026-09-09
- Another mathematician suspects OpenAI trained on his personal chat data — erikphoel · 2026-09-10
Episode 2 · Navier-Stokes dispute raises questions of attribution for aggregated ideas (2026-09-09, 3 posts)
The Navier-Stokes dispute has sparked discussion about AI labs aggregating users' half-formed ideas during training, with researchers advising data-use opt-outs and highlighting attribution as the core unresolved problem.
- Data poisoning research suggests user fragments could shape 'AI discoveries' like Navier–Stokes — MaxDev0 · 2026-09-09
- OpenAI trains on our aggregated thoughts—attribution is the real problem — arjunrajlab · 2026-09-09
- Researcher urges peers to disallow AI training on their data to protect discovery attribution — TuhinChakr · 2026-09-09