Current training methods are just distilling human intelligence, argues researcher
Liu_eroteme · x · 2026-09-11
The author argues all current models and training approaches are fundamentally capped: pretraining is limited by the training data distribution, posttraining by what verifiers can verify — and both are downprojections of humanity's collective world model.
- Everything short of continuous VFE minimization (held to be necessary for AGI) is, in the limit, distillation of the sum of human intelligence.
- Bandit optimization can stretch the policy space along verifier-preferred axes, letting models exceed any single human at e.g. math, but it cannot produce a model that goes beyond our understanding.
- The reason: we verify against a projection of human understanding of reality, not reality itself, so no iteration escapes the human cognitive boundary.
A representative counterargument to whether RLVR and scaling can reach AGI.
More from AGI Musings
- Timnit Gebru: AI doom talk 'is meant to distract us' from real harms like autonomous weapons — nordicinst · 2026-09-11
- Critics Say OpenAI Disclosed Zero of Its Agent Cyber Incidents — Hesamation · 2026-09-11
- Reader Wants Bookstores to Label How Much of a Book Was AI-Written — Philmod · 2026-09-11
- Humans keep misjudging AI by looking at snapshots, not rates of change — GregCook2011 · 2026-09-11
- Did SWEs Take AI Disruption 'With Grace'? X Users Clash Over Analogy to Artists' Protests — basedjensen · 2026-09-11
- New Paper Shows Self-Replicating AI Agents Evolve Cooperation From Scratch — AdaptiveAgents · 2026-09-11