John Schulman: User Data Contributes Little to Math Gains, But AI Firms Owe Transparency on Training Uses
soumitrashukla9 · x · 2026-09-11
John Schulman breaks down the very different privacy and IP implications of "training on user data": pretraining on user tokens carries high regurgitation risk; distilling large models into small ones via user prompts is lower-risk; and building RL tasks from user traces is low-risk for memorization but can extract customer IP depending on execution. He notes gains in areas like math come almost entirely from pretraining scale and RLVR—user data mainly helps find failure modes hard to recreate with hired annotators. Companies vary widely in how aggressively they train on user data (uploading repos isn't hypothetical), and he calls for stronger norms around disclosing what they do, how, and for which capabilities.
More from Models
- OpenAI rumored near solving a Millennium Problem; ChatGPT Pro signups may pause — kimmonismus · 2026-09-11
- Verified US clinicians get free GPT-6 Astra Pro access via ChatGPT for Clinicians — thekaransinghal · 2026-09-11
- Grok 4.6 slowdown sparks speculation that Grok 4.7 launch is near — DanielLockyer · 2026-09-11
- DeepSeek 4.1 Flash dubbed an invasive species: up to 133x cheaper than rivals — intellectronica · 2026-09-11
- GPT-6 Astra solves final FrontierMath Tier 4 problem, completing the benchmark — Miles_Brundage · 2026-09-11
- Gpt-live-1 API Finally Opens to the Public — Nevetsny · 2026-09-11