The decade-old worry that next-token representation would hinder learning turned out unfounded
Aaroth · x · 2026-09-25
@Aaroth reflects on a key shift in how we think about LLMs: a decade ago it was a genuinely interesting computational question whether representing output distributions as next-token distributions would make it hard to learn useful behaviors. It didn't.
- Any text-outputting system can trivially be written as a "next token predictor" — a syntactic fact about probability distributions
- This says nothing about whether brains work similarly to LLMs internally
More from AGI Musings
- Gary Marcus amplifies claim: current agentic frameworks are fully unsafe, need redesign — GaryMarcus · 2026-09-26
- LeCun: Scaling LLMs to AGI Is 'No Way in Hell'; Researcher Pushes Back — aran_nayebi · 2026-09-26
- Philosopher Schwitzgebel: No, We Shouldn't Build AI Guardian Angels — eschwitz · 2026-09-26
- "Normies now simply dislike anything that is AI," observes AI practitioner — BLUECOW009 · 2026-09-26
- Meta and a16z staff claiming AI safety is well-funded draws conflict-of-interest fire — Miles_Brundage · 2026-09-25
- Why AI travel agents can't kill Booking yet — and the case for an agent-native internet — robleclerc · 2026-09-25