Byte-level wrapper lets language models read text letter by letter at near-original cost
bravo_abad · x · 2026-10-08
Most LLMs read text as tokens, so the letters inside a token aren't represented separately—explaining why models write fluently but struggle to rearrange letters, a problem that matters in domains like DNA where one letter can change gene function.
Minixhofer and colleagues retrofit an existing pretrained transformer without rebuilding it: small byte-level layers wrap the model, with an encoder reading byte by byte and grouping bytes into patches that initially mimic the original tokens. The central transformer still processes one unit per patch, keeping cost close to the original, and a decoder writes output byte by byte—giving true character-level manipulation ability.
More from Research
- Eight Knot Shows Autonomous Boat Navigation with ROS at ROSCon JP — 4310sy · 2026-10-08
- ETH's TIDES moves input dependence off the step size so selective SSMs handle irregular time series — ethz · 2026-10-08
- Self-learning agent Gan Jiang hits 96.3% top-1 on XRD phase ID, beating expert skills — Yangtze-ailab · 2026-10-08
- An AI-found proof invents "alien" recursive code — how should we classify its creativity? — burny_tech · 2026-10-08
- Treat AI math proofs like data: why hard-to-read proofs mean more work for humans — inductionheads · 2026-10-08
- DOE and NIH partner with Google DeepMind, Meta on Virtual Biology Initiative to model the cell — DeryaTR_ · 2026-10-08