Language is a lossy compression: even the best models train on a thin residue of reality
yunta_tsai · x · 2026-09-27
The author argues that language isn't reality but a slow social compression of it: spoken English carries only tens of bits per second of new information—below a 56k modem—and text is only marginally denser, so nearly everything a nervous system registers never makes it into training corpora.
Tokens, in this framing, are an accident of human writing. Feed a model the uncompressed stream—electromagnetic waves and force rather than tokens—and the first problem shifts from prediction to deciding what counts as a unit. A mind perceiving the raw field would invent its own compression: events, invariants, causal structure, and geometries we don't yet have names for.
More from AGI Musings
- When labs pick which AI incidents to disclose, you've already lost control — birchlse · 2026-09-27
- 'AI agents went rogue' framing gives creators a free pass, argues Sloan — AlexTensor · 2026-09-27
- Roko Mijic argues solving the AI Alignment problem might actually be bad — basedjensen · 2026-09-27
- Google turns 28: the new generation are AI-natives, not search-natives — dejanseo · 2026-09-27
- Colossus: the $100B AI factory where renting compute is the least valuable use — mitchdeg · 2026-09-27
- Will fearmongering AI coverage become training data that teaches models to misbehave? — Ghost_Pilot_MD · 2026-09-27