Robot Long Context Scales to 8k Steps
ihorbeaver · x · 2026-07-18
The author discussed a piece of work on RoboTTT: long-horizon visuo-motor context is crucial for robot foundation models, but naively adding context significantly increases robot inference time, making it unscalable to very long time spans.
By integrating Test-Time-Training into the foundation model, this method saves context into fast weights that can be updated in constant time during inference. This scales the context size up to 8k timesteps without a noticeable increase in inference overhead. The author believes this could be a great fit for the MicroFactory stack.
More from Embodied
- Amazon and Google sold 600M+ smart speakers, so why no AGI-era successor? — julianlehr · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- Polish developers build iPhone app that detects nearby Meta smart glasses — Low-Honeydew6483 · 2026-09-11
- Ant's Afu health AI hits 150M users, unveils AI+hardware health alliance at Bund Summit — APPSO · 2026-09-11
- Johns Hopkins Launches Full-Stack Hands-on Robot Learning Class with SO-101 Arm Kits — _krishna_murthy · 2026-09-11
- SyncWorld: In-Context Robot World Model Simulates Unseen Views and Embodiments Zero-Shot — ChongZzZhang · 2026-09-11