OpenAI's admission it can't rule out training on anonymized user data sparks research-priority controversy

After OpenAI admitted it "cannot rule out that de-identified product usage data is used for model training," a debate over training data ethics and academic credit quickly spread through the research community. The trigger was user irldanB's complaint that OpenAI could learn from over a year of collaboration with ChatGPT—code records he considered on par with Millennium Prize/Fields Medal level—and, amid rumors he was close to finishing, beat him to the result; he noted Anthropic had a similar case. Israeli NLP researcher Yoav Goldberg pointed out the key distinction: "I accept that a model trained on my data is released publicly" and "I accept an internal model trained on my latest work solving the problem first" are two entirely different things; Boaz Barak added that users can opt out of (de-identified) data being used for training.

Confirmed

Controversy and skepticism

Why it matters

2026-09-09 ~ 2026-09-09 · 9 related posts

Primary sources