Aidan Gomez: labs train on rewritten user data even under ZDR promises
josh_wills · x · 2026-09-09
Cohere co-founder Aidan Gomez says he's heard from employees at both large labs that synthetic data derived from consumer AI tools' production user data is used for training. Users working on 'interesting' tasks (complex math, business, software, bio) are up-weighted. Even ZDR / 'we won't train on you' regimes typically carve out derivative data: only your exact inputs are protected, paraphrased data is fair game.
More from Models
- Astra makes a weird dashboard mistake at just 44% context usage — eigenron · 2026-09-09
- NVIDIA open-sources gold-medal IMO system Nemotron with models, datasets and 200 new problems — kuchaev · 2026-09-09
- Claude asks heavy user for government ID to complete cyber verification — Bulky-Priority6824 · 2026-09-09
- Two years from o1-preview to superhuman math: RL scaling now cracks open research problems — jam3scampbell · 2026-09-09
- Meta's AI assistant outperforms Instinct and ChatGPT Work on a 5-task test, claims reviewer — garrytan · 2026-09-09
- Robot arm self-calibrates with 3 uncalibrated cameras, hits sub-0.2mm accuracy — burny_tech · 2026-09-09