OpenAI's admission it can't rule out training on anonymized user data sparks research-priority controversy
After OpenAI admitted it "cannot rule out that de-identified product usage data is used for model training," a debate over training data ethics and academic credit quickly spread through the research community. The trigger was user irldanB's complaint that OpenAI could learn from over a year of collaboration with ChatGPT—code records he considered on par with Millennium Prize/Fields Medal level—and, amid rumors he was close to finishing, beat him to the result; he noted Anthropic had a similar case. Israeli NLP researcher Yoav Goldberg pointed out the key distinction: "I accept that a model trained on my data is released publicly" and "I accept an internal model trained on my latest work solving the problem first" are two entirely different things; Boaz Barak added that users can opt out of (de-identified) data being used for training.
Confirmed
- OpenAI stated it cannot guarantee user conversations stay out of training data, nor rule out that de-identified product usage data helps improve models.
- One user noted OpenAI began training an updated model on August 28, but discussants considered this unlikely to affect the research conclusions in question, with details involving model variants, fine-tuning, and internal use of models at different training stages.
Controversy and skepticism
- Safety researcher davidmanheim demanded OpenAI tell the truth: either it has no tools to track what data it trained on (unacceptable), or it has the tools but refuses to disclose (worse).
- Developer ctjlewis argued vendors should never promise conversations won't be used for training in the first place, and OpenAI's "excessive honesty" precisely exposes that it can't guarantee it; he asserted that anything anyone has said to GPT or published in the past few months has already gone into the last training round.
- Researcher gleech commented that OpenAI's real situation may be "the data was used automatically, not intentionally, but it's impossible to prove it didn't affect the results—because deep learning is not a science," which he believes is most likely true but extremely damaging for OpenAI.
Why it matters
- The dispute touches on the fundamental trust issue between AI vendors and users/researchers: the untraceability of training data makes "whether someone's work was used to scoop them" nearly impossible to falsify, posing a substantive threat to academic research that relies on open public collaboration.
2026-09-09 ~ 2026-09-09 · 9 related posts
Primary sources
- Yoav Goldberg: Consenting to training data isn't consenting to being raced — yoavgo ·
- OpenAI slammed over training-data opacity: either no traceability or refusing to disclose — davidmanheim ·
- Researcher: OpenAI can't prove training data didn't change results, and 'deep learning isn't a science' — gleech ·
- OpenAI Began Training New Model on Aug 28, Sparking Safety Research Debate — tilmanbayer · 2026-09-09
- [source] OpenAI slammed over training-data opacity: either no traceability or refusing to disclose — davidmanheim · 2026-09-09
- User Claims OpenAI Scooped His Work via Chat Logs, Igniting Data-Training Debate — max_paperclips · 2026-09-09
- Dev on AI labs scooping research: 'Everyone trains on all of it, it's meaningless' — ctjlewis · 2026-09-09
- Dev argues every GPT conversation and paper already went into training, so opt-outs are meaningless — ctjlewis · 2026-09-09
- [source] Yoav Goldberg: Consenting to training data isn't consenting to being raced — yoavgo · 2026-09-09
- Consent to training data isn't consent to AI racing you to your own result, say mathematicians — PMinervini · 2026-09-09
- [source] Researcher: OpenAI can't prove training data didn't change results, and 'deep learning isn't a science' — gleech · 2026-09-09
- Yoav Goldberg: OpenAI's "cannot rule out" user data use is a lawyered admission — MannyKayy · 2026-09-09