Liam Fedus on Scaling Laws: FLOPs are Intelligence, Parameters are Knowledge

hsu_byron · x · 2026-08-20

Liam Fedus retweeted Jie Tang's history of scaling laws, adding context from the 2020 Switch Transformer experiments. They routed tokens to 1 out of 2048 experts, achieving 1.6T total parameters with <3B activated parameters. While this model achieved SOTA on TriviaQA with superior compute efficiency, it failed on reasoning tasks like SuperGLUE. Fedus concluded that the optimal token-per-parameter ratio is task-dependent, reinforcing Noam Shazeer's intuition: FLOPs are intelligence; parameters are knowledge.

Related event: Zhipu's Tang Jie on Scaling Law: GLM-5.3 Gains from Post-training, Not Parameters(9 posts)→

Original post →

More from Research

Research channel →