HLA-WM: training-free hybrid attention gives video world models 12x memory savings
Monash · hf · 2026-10-06
Researchers propose HLA-WM, a training-free hybrid linear-attention framework for long-horizon video world models:
- Motivation: long rollouts need persistent memory, but recurrent linear attention (Gated DeltaNet) suffers severe long-range forgetting as later state updates attenuate distant scenes.
- Method: exploit GDN's affine structure to cache compact chunk-wise transition summaries, retrieve scene-relevant historical chunks via camera geometry, and recompose them into query-specific recurrent states — combining coarse geometry-guided retrieval with fine-grained linear-state computation.
- On the 60-second SANA-WM-Bench, improves all six metrics with no extra training (+0.74 dB PSNR, 28.5% lower rotation error), generalizing to MBench-A.
- At 60-second context, cuts historical-state memory 12x vs full KV caching with at most 1.6% throughput loss.
More from Embodied
- LeCun's AMI Unveils H-JEPA: Hierarchical World Models Lift Visual AntMaze From 18% to 73% — arankomatsuzaki · 2026-10-06
- Tesla appears to be gaining ground on Waymo once data is aligned, scale outside Austin unclear — binarybits · 2026-10-06
- Wayve CEO Alex Kendall: Tesla Proves Self-Driving Demand, But Every Brand Needs to Compete — alexgkendall · 2026-10-06
- Navigating Musculoskeletal Morphospace: A Shape-Shifting Skeleton Demo — zzznah · 2026-10-06
- Soft gains expose sim engine gaps: 16-28mm deviation vs under 1mm in Genesis, robotisist weighs in — chris_j_paxton · 2026-10-06
- Inner Logic's Axel Krieger: surgery is a goldmine for imitation learning — audrow · 2026-10-06