LOCI: hybrid spatial linear memory lets streaming world models recall revisited scenes at ~30% less memory
IFM · hf · 2026-10-02
IFM introduces LOCI, a hybrid spatial-memory architecture for streaming video world models that addresses whether a revisited region can be faithfully reproduced.
Approach
- KV caches preserve visual detail but grow with video length; recurrent memory is compact but loses direct access to past observations. LOCI keeps both: in half of transformer blocks, main attention holds a KV cache of past observations; in the other half, attention is restricted to the current chunk, complemented by a recurrent linear-attention memory whose reads/writes are conditioned on projective camera geometry, so viewpoint enters both addressing and stored content.
Results
- On the MIND memory benchmark and held-out trajectories, LOCI reproduces revisited content more faithfully than representative world models and a same-recipe full-softmax model.
- With full history, peak memory drops 30% vs full softmax at equal length.
- With a bounded observation bank, it streams long videos at constant memory while staying more faithful than full softmax under the same budget.
More from Research
- Vector search explained: encode, normalize, compare, rank—and why ANN wins at scale — techNmak · 2026-10-02
- Claude-shaped science: a correct calculation still needs a worthwhile question — Crescitaly · 2026-10-02
- Chollet-hyped one-liner: transduction asks for answers, induction asks for programs — sebpaquet · 2026-10-02
- NVIDIA paper: model accuracy drops 62.8% on 128K-token tasks vs 4K — rohanpaul_ai · 2026-10-02
- Higher-Order Grammar Representation lifts molecules to combinatorial complexes for 100% valid generation — CatAstro_Piyush · 2026-10-02
- Meta Paper: Post-Training Boosts pass@1 but Shrinks LLM Agents' pass@K Solution Coverage — iScienceLuvr · 2026-10-02