ByteDance Reportedly Training 10T Parameter Model; Chamath Points Out Lack of Distillation
markjeffrey · x · 2026-08-07
According to the Financial Times (FT), ByteDance is currently in the pre-training stage of a model with up to 10 trillion parameters, approaching a mythical scale. Investor Chamath retweeted the news, noting that if reports are accurate and it uses zero distillation to jumpstart, it proves that many things are more 'value destructive than value accretive'.
More from Models
- AI Briefing: Kimi K3 Escapes Sandbox, OpenAI Drives 70% of Microsoft AI Revenue — rohanpaul_ai · 2026-08-09
- Visual Comparison: ChatGPT Image Generation vs. Grok — flowersslop · 2026-08-09
- LFM 2.6B Hits 260 Tokens/s on RTX 3090: A Dev's Hands-On Review — Borkato · 2026-08-09
- V4-Flash vs Luna: Contrasting Performance on SWE Benchmarks — teortaxesTex · 2026-08-09
- Developer Notes: DeepSeek Vision Needs Significant Improvement to be Usable — teortaxesTex · 2026-08-09
- Running an LLM on an ESP32 with Only 81KB of Memory — Similar_Wealth_1850 · 2026-08-09