CMU's CoVeR: Training-Free Token Pruning That Preserves Multi-View 3D Reasoning in VLMs
CarnegieMellonU · hf · 2026-09-09
Carnegie Mellon researchers released CoVeR, a training-free spatial token selector for vision-language models doing multi-view 3D reasoning.
- Prunes multi-view visual tokens while enforcing exact budgets and full scene coverage
- Requires no training and preserves 3D reasoning performance
- Targets the redundancy and cost of multi-view visual token inputs in VLMs
More from Research
- Adding Greek to a Cosmos3 VLA policy: bilingual training helps but lags far behind English — KIEFERSA · 2026-09-09
- Transformers encode a partner's expertise early but only act on it in later layers — Mika Okamoto · 2026-09-09
- Cadence uses a time-series foundation model for error-bounded lossy compression of demand data — Roberto Tacconelli · 2026-09-09
- Meta researcher's training trick: drop a batch proportion instead of per-sample to keep GPU efficiency — TimDarcet · 2026-09-09
- Common Crawl releases experimental graph embeddings for 52.9M web hosts — lhoestq · 2026-09-09
- 3DHarnessBench probes agentic 3D-to-code skills of frontier VLMs — ftm_guney · 2026-09-09