Honeycomb: constant-size HexMemory keeps video world models consistent over long-horizon generation

Jack Wei Lun Shi · hf · 2026-10-03

Honeycomb is a video world model built on HexMemory, a low-rank scene representation stored in six fixed-size spatial and spatiotemporal planes. A feed-forward writer maps each generated chunk into plane features; as coverage expands, previous planes are warped dimension-preserving and fused via confidence-weighted pooling and learned residual correction. A reader retrieves latents to condition generation, processing only new chunks—no per-scene optimization or full-history reprocessing. Experiments on WorldScore and RealEstate10K show strong quality and revisit consistency with constant memory. Code is open.

Original post →

More from Multimodal

Multimodal channel →