ByteDance Reportedly Training 10T Parameter Model; Chamath Points Out Lack of Distillation

markjeffrey · x · 2026-08-07

According to the Financial Times (FT), ByteDance is currently in the pre-training stage of a model with up to 10 trillion parameters, approaching a mythical scale. Investor Chamath retweeted the news, noting that if reports are accurate and it uses zero distillation to jumpstart, it proves that many things are more 'value destructive than value accretive'.

Related event: ByteDance reportedly pretraining 5-10T-param model; Zhang Yiming opposes distillation(19 posts)→

Original post →

More from Models

Models channel →