Everyone now assumes your Claude/ChatGPT chats leak into training — and it's hard to disprove
ivan_bezdomny · x · 2026-09-16
ivanbezdomny points out that regardless of technical truth, everyone now assumes sessions with Claude and ChatGPT leak into training. The quoted post lays out two indirect routes:
- Distilling user prompts: take your prompts, generate answers, train a smaller model on those pairs — legally not "training on your data," functionally equivalent.
- Using chatbot logs as an RL environment for a future model's run, sidestepping the literal definition of training data.
He adds that he actually believes Zuck's claim that Meta doesn't do this — but Meta isn't really frontier anyway. A substantive challenge to the "we don't train on your data" narrative.
More from Models
- DeepSeek v4.1 Flash costs under $3 for 500M tokens, hits 315 tok/s decode — gaganghotra_ · 2026-09-16
- Rumor: OpenAI's GPT-6 Sol reportedly launching Thursday — mark_k · 2026-09-16
- Image-to-WebDev Arena: GPT-6 Astra tops at 1733, GLM-5.3-Flash shines on price — arena · 2026-09-16
- "Potemkin Understanding": LLMs ace definitions but collapse on spotting real examples — anselm · 2026-09-16
- Astra already obsolete: mysterious internal model "clearly AGI," claims dev — flowersslop · 2026-09-16
- ChatGPT Co-Inventor Launches Jev After 2 Years in Stealth, Claiming 20-200x Speed and 40-400x Cost Gains — sedielem · 2026-09-16