PLANET lifts multi-object tracking into 3D scene geometry, hits SOTA on DanceTrack, SportsMOT & BFT
lealtaixe · x · 2026-09-07
Researchers including Laura Leal-Taixé introduce PLANET (arXiv:2609.00924), a multi-object tracker that moves beyond the image plane.
Key ideas:
- Monocular video is a 2D projection of 3D scenes; trackers relying only on image-plane appearance and geometry inherit depth ambiguity
- Existing 2D tracking datasets are lifted into 3D as an enabling step
- Reconstructed 3D scene geometry is embedded into query features and positional encodings, plus an auxiliary 3D location prediction task
- A dual-resolution temporal memory preserves evidence across longer gaps
PLANET achieves state-of-the-art results on DanceTrack, SportsMOT, and BFT.
More from Research
- Toby Ord: a photo can convey 512 bits, but not freely chosen ones — tobyordoxford · 2026-09-07
- Toby Ord on the two senses of how many bits a photograph can communicate — tobyordoxford · 2026-09-07
- DeepMind-Princeton paper shows LLMs causally use confidence to decide whether to answer — GoogleDeepMind · 2026-09-07
- The attention triangle: diagnosing cross-modal semantic leakage in audio-video diffusion — tau · 2026-09-07
- Yandex researchers propose KV cache as an agent runtime, demo Qwen3.8 playing DOOM interactively — _puhsu · 2026-09-07
- ECCV 2026 workshop on deep learning era SfM set for Sept 9 with 3 speakers and 8 papers — ducha_aiki · 2026-09-07