Modern LLMs Compress English Text to Under 1 Bit Per Character
docmilanfar · x · 2026-08-27
Reflecting on Shannon's estimate 75 years ago that English is 75% redundant, the entropy is no more than 1 bit per letter when considering long-range context. Only recently have text encoders achieved this rate. In fact, modern LLMs predict the next character so accurately that they routinely compress English text to under 1 bit per character.
More from Research
- GPT-5.6 Builds New Kernel, Achieving 9.7x Speedup on TPU — HuaxiuYaoML · 2026-08-27
- RSI-Exam Benchmark Launches to Test AI Recursive Self-Improvement — HuaxiuYaoML · 2026-08-27
- Gordian Screens 1,327 Targets In Vivo, Accelerating Drug Discovery — juanbenet · 2026-08-27
- Discussion on Multi-Agent Reward Schemes and Convergence — jessi_cata · 2026-08-27
- Why scaling LLMs won't lead to real agency: A 3-tier Embodied AI architecture — Far-Start-1789 · 2026-08-27
- Terence Tao on Human-AI Complementarity: AI Excavates, Humans Recognize — bennash · 2026-08-27