Distilling Transformers into Recurrent Ones: Linear-Time Robot Memory That Matches Full History
chriswolfvision · x · 2026-10-06
A new arXiv paper proposes distilling a full-history Transformer into a recurrent Transformer for long-horizon streaming vision and robotics, where storing observation history is impractical.
- Claim: the performance gap isn't architectural—recurrent models face a much harder learning problem, deciding at each step what to keep in fixed-size memory.
- Method: a teacher explicitly compresses observation history into a fixed-size bottleneck representation, which directly supervises the student's memory, aligning the two compression mechanisms.
- Result: a recurrent latent robotic memory with linear-time complexity that approaches full-history Transformer performance.
More from Embodied
- coprod Launches co:system: One EtherCAT Bus and Sub-ms Real-Time Control for Whole Robots — FlolightC · 2026-10-06
- heypcb maps 11,000+ open-source hardware boards because GitHub wasn't built for hardware — Stefania_druga · 2026-10-06
- Robotics researcher slams microfactory demo as 'too perfect' sim missing real-world contact dynamics — ChongZzZhang · 2026-10-06
- YC robot data startup wins customers in a commoditized market by out-supplying rivals — paigeinsf · 2026-10-06
- Flashing Telink firmware via Flipper Zero to hack Tuya Zigbee devices — ngxson · 2026-10-06
- Norway to propose temporary ban on AI glasses in some public places — talkingatoms · 2026-10-06