Labs reportedly turn proprietary chain-of-thought into synthetic training data; TOS only covers user-visible IO
erikphoel · x · 2026-09-10
erikphoel quotes cullend's claim and asks for verification: both OpenAI and Anthropic's terms of service only promise not to train on user-visible inputs/outputs after users opt out of data training — internal reasoning is not covered.
The more striking claim: for the past 20 months, turning reasoning/chain-of-thought (which labs treat as proprietary and don't show users) into synthetic training data has been the primary way labs generate new data. If true, even users who opt out could have long reasoning traces feeding training pipelines. Unverified.
More from Models
- DeepSeek's answer to surging demand: make its model cheaper and faster — yacineMTB · 2026-09-10
- Prelim assessment: GLM-4.1 behaves similar to V4, Zhipu's post-training called more advanced — menhguin · 2026-09-10
- Huge share of post-2022 web data is AI content mislabeled as human-written — menhguin · 2026-09-10
- DeepSeek V4.1 Hailed as the 'First Gamer Model' After Blowing Away an FPS-Generation Test — teortaxesTex · 2026-09-10
- Bug Hunt Bench grades frontier models on 105 real bugs; DeepSeek-V4.1-Flash lands 24/105 for $1.80 — PawelHuryn · 2026-09-10
- RSI is here, just disaggregated: DeepSeek using LLMs to design algorithms — teortaxesTex · 2026-09-10