Meta paper shows byte-level models beat tokenized ones given enough training compute
alex_verem · x · 2026-10-10
A new Meta FAIR / University of Washington paper challenges the field's reliance on tokenizers: a small model reading raw bytes (vocab 256) can surpass a tokenized model of similar size given enough training.
- Distilled from a Llama 3-8B teacher, the byte student has fewer total parameters (1.28B vs 1.81B) yet wins
- It matches the token model's accuracy with 1/6 of the training data and cuts teacher data storage to 1/5
- Scaling laws project 4% higher accuracy across six benchmarks, beating Llama 3.2-1B and Gemma-3-1B-pt
- The trick: a marker every 4.5 bytes signals token boundaries, costing 31% more compute per chunk
Takeaway: the tokenizer may be a habit, not a necessity.
More from Models
- "Fake videos are on another level now": researcher shares hyper-realistic AI video — rohanpaul_ai · 2026-10-10
- Leak claims major model drops next week, dismissing slowdown talk — iruletheworldmo · 2026-10-10
- Haiku flubs trading basics: can't tell bids from offers, user reports — arthurcolle · 2026-10-10
- Step 5 Preview builds a working dashboard with zero code changes, demo shows — StepFun_ai · 2026-10-10
- HiDream-O1-Video-1.0 debuts #6 on image-to-video leaderboard at $5.80/min — ArtificialAnlys · 2026-10-10
- Blogger Speculates Anthropic Had Opus 5.5-Level Intelligence Internally Six Months Ago — yihui_indie · 2026-10-10