John Schulman: User Data Contributes Little to Math Gains, But AI Firms Owe Transparency on Training Uses

soumitrashukla9 · x · 2026-09-11

John Schulman breaks down the very different privacy and IP implications of "training on user data": pretraining on user tokens carries high regurgitation risk; distilling large models into small ones via user prompts is lower-risk; and building RL tasks from user traces is low-risk for memorization but can extract customer IP depending on execution. He notes gains in areas like math come almost entirely from pretraining scale and RLVR—user data mainly helps find failure modes hard to recreate with hired annotators. Companies vary widely in how aggressively they train on user data (uploading repos isn't hypothetical), and he calls for stronger norms around disclosing what they do, how, and for which capabilities.

Original post →

More from Models

Models channel →