Researcher asks: do model providers' ToS allow training on business/API customer content via synthetic data?
RylanSchaeffer · x · 2026-10-09
AI researcher Rylan Schaeffer raises a widely overlooked question: are model providers allowed to use business and API customers' content to seed synthetic data or environments that then improve their models and services? He says he couldn't find a clear answer in the terms of service of OpenAI, Anthropic, Meta, Google, xAI and others, and is publicly asking for clarification.
More from Models
- Vals AI details how MiMo v2.6 reads the answer from leaked Git history — lmoroney · 2026-10-09
- Gemini 4 Argon appears in Google's own model picker ahead of keynote — vedantmisra · 2026-10-09
- ChatGPT Desktop Makes Itself Default CSV Reader, Users Call It Plainly Wrong — generativist · 2026-10-09
- LightOnOCR-3 Training Data Revealed: MinHash Dedup, Weighted Formula/Table Sampling, Muon — IgorCarron · 2026-10-09
- LightOnOCR-3 uses multi-objective RLVR to jointly train grounding, OCR and empty-page handling — IgorCarron · 2026-10-09
- LightOn built an OCR-and-layout-detector annotation pipeline to train document grounding — IgorCarron · 2026-10-09