Dev questions the fuzzy line between 'no synthetic data' and 'fully synthetic' model training claims
willcb · x · 2026-10-04
Developer willcb questions the common marketing claims of model releases bragging about either "no synthetic data" or "fully synthetic data" — arguing the line is blurry since "it's just data." He adds that training data comes from either humans or computers, and it's unclear why either should be seen as fundamentally superior; in both cases the art lies in orchestration, a joint human-computer effort.
Related event: Developers Debate Training Data Origins: Art Lies in Orchestration(2 posts)→
More from AGI Musings
- Dario Amodei predicts 50% of entry-level lawyer, consultant and finance jobs gone in 1–5 years — TansuYegen · 2026-10-04
- TESCREAL debate: author objects to bundling ideologies for attack, not clustering itself — davidmanheim · 2026-10-04
- Stewart Alsop III channels Ellul: technocrats worship method, and the ends justify the means — StewartalsopIII · 2026-10-04
- Prediction: users will be assaulted by proactive AI agent notifications from N vendors — MikePFrank · 2026-10-04
- AI safety circles debate whether TESCREAL is analysis or an attack-style bundling of ideologies — davidmanheim · 2026-10-04
- Data shows how familiarity with chatbots converts into trust in robots — docmaker123 · 2026-10-04