Researcher critiques scaling experiment: identical architectures and multiple-choice evals skew results

burny_tech · x · 2026-09-09

Elie Bakouch raises two technical critiques of @dwarkeshsp's scaling experiment: the tested architectures are all dense full-attention transformers differing only in swiglu, qk-norm and rope values, with no MoE, so conclusions about architecture's (in)significance are limited; and the evals appear to be multiple-choice QA instead of completion/PPL-based evals, which are easily benchmaxed by including QA data that newer datasets filter out. He still finds the experiment interesting and calls for public checkpoints for further study.

Original post →

More from Research

Research channel →