Debate: is pretraining size scaling dead, or just architecture scaling conflated?
samsja19 · x · 2026-09-09
- Responding to Sarah Hooker's claim that pretraining model-size scaling is dead because transformers are saturated, the author questions whether people are still scaling model size and whether the model axis in question measures architecture rather than size.
- Highlights ongoing confusion in scaling debates between scaling model size and scaling architectures.
More from Models
- Astra tested on ~30 obscure puzzle games: ARC-AGI-3's ~99% may undersell it — burny_tech · 2026-09-09
- OpenAI claims it "solved Navier-Stokes"; Pedro Domingos calls claim ignorant or dishonest — jonathanberte · 2026-09-09
- Meta's Muse Spark 1.3 Max hits 67.9% on CursorBench at $1.31/task, 4.3x cheaper than GPT-5.6 Sol Max — shuyanzh36 · 2026-09-09
- Leak: OpenAI internal model 'bel' reportedly solves ~2.5x more math problems than astra — imjustnewatai · 2026-09-09
- GPT-6-Astra Lands in Loop, the Debugging Tool Built on Codex — nikunjhanda · 2026-09-09
- Founder praises GPT Astra for faster reasoning and sharper interpersonal guidance — morganlai · 2026-09-09