Report: ByteDance Pre-Training 10-Trillion Parameter AI Model, Far Exceeding Kimi

eyishazyer · x · 2026-08-07

According to a report by the Financial Times (FT), ByteDance is currently pre-training an AI model with up to 10 trillion (10T) parameters, vastly exceeding the scale of the recently popular Kimi K3 (2.8T parameters).

The report also notes that ByteDance has avoided distilling rival models for over a year, preferring independent development. If this reported run succeeds, it will demonstrate ByteDance's ability to execute frontier-scale pre-training without leaning on a competitor's model as a teacher.

Related event: ByteDance reportedly pretraining 5-10T-param model; Zhang Yiming opposes distillation(19 posts)→

Original post →

More from Models

Models channel →