PersistBench (NeurIPS Spotlight): 4D foundation models can see but not remember

weichiuma · x · 2026-10-07

PersistBench, from a Cornell team, earned a Spotlight at NeurIPS 2026 (Evaluations & Datasets Track). It asks: 4D foundation models can perceive and reconstruct dynamic environments, but can they remember what they perceived, as humans do? Existing benchmarks rely on pixel-level metrics and lack ground truth once objects leave the field of view.

Method: Uses 360° videos as omniscient ground truth, sidestepping the challenge of knowing what happens to objects off-camera, and evaluates three object-centric aspects of visual memory: object permanence, motion continuity, and appearance preservation.

Findings: Across 12 models from three families, evaluation reveals a substantial gap between synthesizing visible content and remembering objects beyond the input view. Project page, leaderboard, code and dataset are all open for submissions.

Original post →

More from Research

Research channel →