The scaling lesson: dismissed architectures may just be data-starved
Silver-Champion-4846 · reddit · 2026-09-21
A Reddit discussion argues that GPT-1 to GPT-6 share the same decoder-only architecture, differing only in data scale, quality, SFT and RL — so early capabilities shouldn't be read as an architecture's ceiling. The author speculates that newer models like Jev and Laya hint encoder-only models scaled up with internet-scale data could surprise us too, asking which dismissed architectures were merely data-starved.
More from AGI Musings
- Devs clash over whether beginners should still learn to code as AI automates code review — mgill25 · 2026-09-21
- OpenAI's Noam Brown: aligned AI workers could hand the edge to incumbents over startups — victor_explore · 2026-09-21
- If your p(doom) > 0, why are you accelerating frontier lab research? — _arohan_ · 2026-09-21
- AI assistants don't remove decisions, they multiply them — and nobody wants that — menhguin · 2026-09-21
- Daniel Mac: It's amazing we can even debate whether current AI is AGI — daniel_mac8 · 2026-09-21
- Reddit hot take: AI safety regulation push is a cartel raising rivals' costs — crua9 · 2026-09-21