DeltaWorld: One Delta Token Per Frame Cuts Video World Model Tokens 1024x
gabriberton · x · 2026-09-25
Tommie Kerssies and colleagues introduce DeltaWorld, a generative world model built on DeltaTok, a tokenizer that encodes the vision-foundation-model feature difference between consecutive frames into a single continuous "delta" token.
Key points:
- Reduces video from a 3D spatio-temporal representation to a 1D temporal sequence — e.g. a 1,024x token reduction on 512x512 frames
- Compact representation makes multi-hypothesis training tractable: many futures generated in parallel, only the best supervised
- At inference, yields diverse plausible future predictions in a single forward pass
- Outperforms on dense forecasting tasks versus existing generative world models, at far lower compute
More from Research
- DeCAF flow-map framework speeds protein-ligand cofolding 5x, earns NeurIPS Spotlight — rishabh16_ · 2026-09-25
- WROP: 150 cognitive tasks train object permanence into video world models — Haotian Zhang · 2026-09-25
- AEWM: editing agent task states beats simulating tool responses — RUC · 2026-09-25
- Qwen-Planner-Agent: a closed-loop AI-for-AI framework for mobile planner agents — Tongyi-MAI · 2026-09-25
- SF Systems Meetup talk explores bringing programming languages beyond text — ShadajL · 2026-09-25
- NVIDIA's Nemotron-Cascade accepted as NeurIPS 2026 oral for cascaded domain-wise RL — ctnzr · 2026-09-25