Controlled Study Finds No Encoding Dominates: Pixels, Bytes and Tokens Each Win on Different Tasks
delliott · x · 2026-10-06
A new paper, Rate-Utility Frontiers for Language Encodings by Ingo Ziegler, Martin Krebs and Desmond Elliott, compares subword tokens, raw bytes, and rendered pixels under carefully controlled conditions.
- Using verified parallel sentences across 13 languages and 5 scripts, the authors sweep a shared bottleneck to trace rate-utility frontiers, separating sequence length, latent capacity, and task-relevant information that survives compression.
- Three utilities are evaluated: surface form preservation, cross-lingual sentence alignment, and topic classification.
- No encoding dominates: pixels best preserve surface form, bytes best preserve cross-lingual alignment (especially in same-script multilingual settings), and tokens best support topic prediction. Performance is not explained by sequence length alone.
The same thread also highlights a second paper showing that open-weight VLMs do not appear to revise their predictions during CoT reasoning.
More from Models
- OpenAI shares revenue with popular plugins, model picker may go away, says ChatGPT head — gaganghotra_ · 2026-10-06
- Mustafa Suleyman once claimed hallucinations would be 'largely eliminated' by 2025 — PMinervini · 2026-10-06
- GPT-6 Astra bank-reset abuse blamed for ChatGPT Pro 20x weekly quota halving — teortaxesTex · 2026-10-06
- Laya: Open-Source Model Returns Calibrated Probabilities in 33ms Across 100+ Languages — victormustar · 2026-10-06
- Lab Trains Qwen3-32B to Introspect: Faithful Self-Report Emerges Late, With a Measurable Neural Footprint — davidbau · 2026-10-06
- Claude suddenly refuses to edit Terraform tfvars files after Anthropic guardrail update — basedjensen · 2026-10-06