Memorizon trains streaming world models beyond context window with only 12% step-time overhead

MBZUAI-IFM · hf · 2026-10-02

MBZUAI-IFM's Memorizon decouples supervision span from attention span for streaming world models: scored chunks retrieve top-K latents via camera co-visibility into a shared bank bounded by kK, so long-span training stays cheap. Extending spans from 100s to 400s adds just 12% step time, while reaching a location's first visit boosts revisit consistency by 24-30% over a sliding-window baseline. Filling the bank from another episode cuts revisit correlation by 83%, confirming the model actually uses retrieved latents.

Original post →

More from Research

Research channel →