Labs reportedly turn proprietary chain-of-thought into synthetic training data; TOS only covers user-visible IO

erikphoel · x · 2026-09-10

erikphoel quotes cullend's claim and asks for verification: both OpenAI and Anthropic's terms of service only promise not to train on user-visible inputs/outputs after users opt out of data training — internal reasoning is not covered.

The more striking claim: for the past 20 months, turning reasoning/chain-of-thought (which labs treat as proprietary and don't show users) into synthetic training data has been the primary way labs generate new data. If true, even users who opt out could have long reasoning traces feeding training pipelines. Unverified.

Original post →

More from Models

Models channel →