Memorizon: training world models beyond their context window with minute-long video

ZhitingHu · x · 2026-10-06

A new research work, Memorizon, tackles a core weakness of world models: trained on 10-second clips, they rarely see both the first visit to a place and a return, so they never learn long-horizon consistency. Memorizon trains on minutes-long video while each chunk attends only to a small bank of retrieved frames, keeping cost bounded.

Original post →

More from Research

Research channel →