FULL STORY

ByteDance Reportedly Pretraining 5-10 Trillion Parameter AI Model

ByteDance is reportedly developing a massive 5-10 trillion parameter AI model to catch up with US AI leaders, with founder Zhang Yiming explicitly opposing distillation.

2026-08-07 ~ 2026-08-09 · 2 episodes · 21 posts

Episode 1 · ByteDance reportedly pretraining 5-10T-param model; Zhang Yiming opposes distillation (2026-08-07, 19 posts)

According to reports from the Financial Times and other media, ByteDance is advancing early pretraining of a massive model with 5 trillion to 10 trillion (5T-10T) parameters, expected to be ready by the end of this year. Founder Zhang Yiming has explicitly opposed distilling Western models, requiring the team to insist on independent R&D. If the rumors are true, this would be the largest known model in China, marking a new level in the large model arms race.

Confirmed

  • According to reports from the Financial Times and other media, ByteDance is pretraining a 5T-10T parameter model, currently in early stages, expected to be completed by year-end.
  • ByteDance founder Zhang Yiming explicitly stated at an internal meeting that he firmly opposes distilling Western models, requiring no shortcuts and insisting on independent R&D of underlying capabilities.
  • The project is led by ByteDance's Seed team, which has grown to about 2,000 people globally. It is reportedly led by former Google DeepMind employee Wu Yonghui and Seed Foundation head Xiang Liang.

Unconfirmed

  • The final parameter scale of the model is still uncertain, with rumors ranging from 5T to 10T; the final size has not been decided.
  • The specific technical route and actual deployment method of the model remain unclear.

Why it matters

  • If true, this would be the largest known model in China, several times the size of Kimi K3 (2.8T), and directly comparable to Anthropic's rumored Mythos model with about 8T parameters.
  • Regarding the feasibility of deploying a 10T model, netizen @zephyrz9 and researcher Alex Suchenzang have raised doubts, arguing that directly serving a 10T model online is unrealistic. They speculate ByteDance will likely need to use techniques like model distillation to compress it into smaller, more efficient versions for practical use.
  • Industry insider @zijingwu further pointed out that pretraining such a giant model actually requires only about 30,000 GPUs; the real compute challenge lies in inference cost, so vendors will likely end up offering distilled versions.
  • Investor Chamath noted that if the report is true and the model was started without any distillation, it would prove many things are "value destruction" rather than "value creation."

Episode 2 · ByteDance Reportedly Training Massive AI Model to Catch Up with US Labs (2026-08-09, 2 posts)

ByteDance is reportedly training a massive new AI model aiming to rival Anthropic's most advanced systems. This move highlights a broader effort by Chinese AI labs to accelerate development and close the technological gap with top US labs.