OpenAI slammed over training-data opacity: either no traceability or refusing to disclose
davidmanheim · x · 2026-09-09
- Safety researcher davidmanheim presses OpenAI on its statement that it 'cannot rule out' de-identified usage data helped improve its models: why won't it say what the ground truth is?
- His dilemma: either OpenAI lacks tooling to trace what it trains on (unacceptable), or it has it but won't disclose (damning).
- He also notes the newer model OpenAI began training on Aug 28 could have ingested the data, and argues the influence is more plausible than skeptics claim since training continued through a later date on the model in question.
More from Companies & People
- Stealth mode is dying: founders stop sharing ideas because they're too easy to copy — deliprao · 2026-09-09
- StepFun heads to AGNTCon + MCPCon Japan (Sept 10-11, Tokyo) at booth T6 — StepFun_ai · 2026-09-09
- DeepSeek's official X account called out for ignoring its own price cuts and releases — teortaxesTex · 2026-09-09
- Sarvam AI and IDFC FIRST Bank Launch Joint Lab to Build a Self-Improving Bank — itsOmSarraf_ · 2026-09-09
- Distillation as IP theft vector: legal language frames a legitimate training technique as a weapon for rivals — Xianbao_QIAN · 2026-09-09
- DeepSeek Opens ~150 Engineering Roles, Zero Research Positions as Race Shifts to Systems — pstAsiatech · 2026-09-09