Robot Long Context Scales to 8k Steps

ihorbeaver · x · 2026-07-18

The author discussed a piece of work on RoboTTT: long-horizon visuo-motor context is crucial for robot foundation models, but naively adding context significantly increases robot inference time, making it unscalable to very long time spans.

By integrating Test-Time-Training into the foundation model, this method saves context into fast weights that can be updated in constant time during inference. This scales the context size up to 8k timesteps without a noticeable increase in inference overhead. The author believes this could be a great fit for the MicroFactory stack.

Original post →

More from Embodied

Embodied channel →