If US and Chinese labs buy training data from the same vendors, what makes models different?
biotechplays · reddit · 2026-09-12
A Reddit discussion asks: if AI labs buy training data from the same vendors, what actually differentiates their models — coding strength, writing, instruction following? Citing a Forbes report that US data vendors also supply Chinese labs, the poster wonders how much the underlying training material overlaps across labs. The thread debates how much capability comes from the data itself versus what each lab does with it — curation, mixture, and post-training.
More from Models
- GPT-6 Astra review: stunning at 3D games and computer use, still not a daily driver — petergyang · 2026-09-12
- First quantitative evidence: Claude and GPT now use GUIs as well as APIs — ysu_nlp · 2026-09-12
- DeepSeek 4.1 Flash ignores instructions — and that's by design, says fan — teortaxesTex · 2026-09-12
- Perplexity trusts GPT-6 Astra with end-to-end production systems — OpenAI News · 2026-09-12
- GLM-5.3 Scores 28.8% on New RealSWE Benchmark, Closing In on GPT-6 Astra — zainhas · 2026-09-12
- Open models flop on Terminal-Bench Science: best scores just 4/70 — teortaxesTex · 2026-09-12