DEER-3D fixes 3D-LLM grounding biases with error-driven counterfactual scene editing

YonatanBitton · x · 2026-09-10

An ECCV paper introduces DEER-3D, an error-driven scene editing framework improving 3D visual grounding in 3D-LLMs. It runs a structured 'Decompose, Diagnose, Edit, Retrain' loop: identifying predicate-level grounding failures (attribute or spatial relations), applying minimal predicate-aligned scene edits like recoloring or repositioning, and constructing targeted counterfactual question-answer pairs for retraining — avoiding costly scene reconstruction or large-scale 3D data collection. Evaluations across multiple 3D grounding and scene understanding benchmarks show consistent improvements in spatial and attribute grounding.

Original post →

More from Multimodal

Multimodal channel →