Depth-to-width ratio is the parameter that matters, industry has converged, says Stanford NLP
stanfordnlp · x · 2026-09-05
Stanford NLP weighed in on a CS336 discussion about model scaling, pointing out that the depth-to-width ratio (plus MoE) is what matters most — and that the industry has more or less already converged on this parameter, with most models fairly close to each other on it.
More from Research
- Xbench turns Twitter into an AI eval, tracking real sentiment and model switches — thedealdirector · 2026-09-06
- Principia: A Benchmark Testing Whether Video Models Grasp Physics via Pendulum Dependencies — anand_bhattad · 2026-09-06
- Jerry Liu: The Other Half of RLMs Is Programmatic Recursion, aka 'Dynamic Workflows' — lateinteraction · 2026-09-06
- Blogger pegs 30% odds OpenAI already solved Navier-Stokes, 50% partial progress — scaling01 · 2026-09-05
- LLMs write locally coherent but globally incoherent quests: an MMO writer's thousands-of-quests problem — HLCYSWAP · 2026-09-05
- Why Agents Last Exam Scores Jumped: Computer Use, Bigger Models, Diverse RL Envs — dejavucoder · 2026-09-05