Why AI doesn't actually read words: from BPE subwords to byte-level models like BLT and H-Net
jbhuang0604 · x · 2026-10-09
- JB Huang (Tencent chief AI scientist) released an explainer video on how LLMs actually represent text — they don't read words the way humans do.
- The video traces the technical lineage from words and BPE-style subword tokens to byte-level models including ByT5, CANINE, MegaByte, BLT, and H-Net, covering the evolution and trade-offs of tokenization schemes.
More from Research
- Ex-self-driving ML engineer writes long-form on the practice of semi-supervision — Visual_Ability · 2026-10-09
- RLVR misses 'all minimal correct answers' problems; new credit assignment doubles finds — thoma_gu · 2026-10-09
- Debunked: AI did not solve the Millennium Prize Navier-Stokes problem — gerardsans · 2026-10-09
- New open-source 3JSBench evaluates LLMs on generating coherent Three.js 3D assets — ycombinator · 2026-10-09
- Moonworks' Lunara: Sub-10B Diffusion Mixture Transformer Tops Aesthetic and Human Blind Evaluations — paper-crow · 2026-10-09
- Study: Local domains take 41.4% of AI citations in Brazil, 38.3% in UK across ChatGPT and Gemini — gaganghotra_ · 2026-10-09