VideoLoop rewrites bounded working memory, hits 88.3% on VideoMME long video
Jinfa Huang · hf · 2026-09-30
Long-video agents suffer semantic thrashing: append-only working memory grows unboundedly, attention to key evidence collapses, and the agent loses track of what it already found.
- Structural argument: append-only memory can add new evidence but cannot remove accumulated noise or stop ordered context growth without a rewrite operator
- VideoLoop couples two loops: an outer reasoning loop over the video, and an inner loop that retrieves artifacts from an unbounded filesystem of past observations and rewrites bounded working memory each step
- Plug-and-play gains of 4.2 pts on average across four LVLM backbones on VideoMME (long)
- On the hardest quarter of VideoMME (long), a blind judge reading only context answers 81.1% vs 60.9% for the append-only agent
- With Gemini 3.1 Pro: 88.3% VideoMME (long), 88.8% VideoMMMU, 80.9% LongVideoBench (long)
More from Multimodal
- AI-reanimated Greta Garbo stars in SKF ball-bearing ad, panned as bland — nordicinst · 2026-09-30
- Dev predicts a universal DSL for video data will enable highly controllable world-simulation diffusion models — zeeshanp_ · 2026-09-30
- MiniMax H3 at max settings takes 26 min and 62GB VRAM per 15s clip on RTX Pro 6000 — Realistic-Fennel-190 · 2026-09-30
- I used Codex to make a 15-second animation — and still had to give it editing notes — naridubs · 2026-09-30
- Flux 3 nails four-way split-screen images of one event from four angles — umesh_ai · 2026-09-30
- MageTrail 2.8B booru finetune costs $593 so far, hits limits of 41k-image dataset — Turbulent-Bass-649 · 2026-09-30