MemoBench: Top 10 Video Models Fail to Remember Object States During Occlusion

jiqizhixin · x · 2026-07-06

Researchers from Harvard, MIT, and Google introduced MemoBench to evaluate whether video generation models can correctly update an object's state after it disappears from view (e.g., when the camera pans away from melting ice or pouring powder and returns). Evaluations of 10 top-tier video models revealed that none could reliably maintain memory across occluded frames, highlighting a major open challenge in building current world models.

Related event: MemoBench Reveals Top Video Models Fail at Object Permanence(2 posts)→

Original post →

More from Multimodal

Multimodal channel →