Scaling01 on Astra: no AGI vibes, lacks taste, doubts 10T params, guesses 6-8T
scaling01 · x · 2026-09-07
- KOL scaling01, after burning through his Pro limits on Astra, says it doesn't feel like AGI: the model lacks conviction and taste, wasting time and tokens.
- He discounts rumors of 10T+ parameters — if true, scaling laws would be "cooked" — and guesses 6-8T instead.
- He also mocks the idea of a 100-day pre-training run, arguing it would waste 100 days of algorithmic progress, and suggests short iterations may be the base case for GPT-6. His verdict: Astra is just a slightly larger code monkey — the best available, but not AGI.
Related event: KOL Tests Astra, Sees No AGI; Speculates GPT-6 Uses Up to 5x Compute(3 posts)→
More from AGI Musings
- Geoffrey Hinton admits he was wrong about AI replacing radiologists — and explains why — Afinetheorem · 2026-09-07
- Open Offices Were Onto Something — But They Need Mature Ambient Compute to Work — curious_vii · 2026-09-07
- Autoregressive models nailed sheet music back in 2019 — LLMs just haven't been fed it — fly_ght · 2026-09-07
- Ex-xAI researcher Ethan He: finding the right axis to scale matters more than raw compute — ricklamers · 2026-09-07
- Seth Lazar: Social sciences must self-critique before asking labs for funding — sethlazar · 2026-09-07
- As AI drives creation costs to zero, the 'handmade' premium on creative work erodes — joonasvirtanen · 2026-09-07