Language is a lossy compression: even the best models train on a thin residue of reality

yunta_tsai · x · 2026-09-27

The author argues that language isn't reality but a slow social compression of it: spoken English carries only tens of bits per second of new information—below a 56k modem—and text is only marginally denser, so nearly everything a nervous system registers never makes it into training corpora.

Tokens, in this framing, are an accident of human writing. Feed a model the uncompressed stream—electromagnetic waves and force rather than tokens—and the first problem shifts from prediction to deciding what counts as a unit. A mind perceiving the raw field would invent its own compression: events, invariants, causal structure, and geometries we don't yet have names for.

Original post →

More from AGI Musings

AGI Musings channel →