Researcher critiques scaling experiment: identical architectures and multiple-choice evals skew results
burny_tech · x · 2026-09-09
Elie Bakouch raises two technical critiques of @dwarkeshsp's scaling experiment: the tested architectures are all dense full-attention transformers differing only in swiglu, qk-norm and rope values, with no MoE, so conclusions about architecture's (in)significance are limited; and the evals appear to be multiple-choice QA instead of completion/PPL-based evals, which are easily benchmaxed by including QA data that newer datasets filter out. He still finds the experiment interesting and calls for public checkpoints for further study.
More from Research
- Self-trained model switches to weekly Sunday weight updates for richer reward gradients — cephaloform · 2026-09-09
- Preregistered study finds both writing skill and CS background predict vibe-coding performance — dair_ai · 2026-09-09
- Frontier AI now doing advanced math — and getting into academic priority disputes — anshulkundaje · 2026-09-09
- Dev uses Claude to visualize OpenAI's finite-time blowup construction for 3D Euler — CatAstro_Piyush · 2026-09-09
- SimpleMemVLA feeds full video history to a VLM, beating dedicated memory modules for long-horizon robot manipulation — openbmb · 2026-09-09
- Elicit modeling exercise estimates indoor/outdoor living costs ~2 years of lifespan — elicitorg · 2026-09-09