VIGA agent rebuilds images into editable Blender scenes via multimodal inverse-graphics loop

Michael_J_Black · x · 2026-09-12

Researchers from UC Berkeley, CMU, and Max Planck present VIGA (Vision-as-Inverse-Graphics Agent) at ECCV: a multimodal agent that reconstructs input images as editable scene programs in Blender via an analysis-by-synthesis loop with interleaved multimodal reasoning and evolving contextual memory. It can build scenes from primitives or leverage tools like Meshy and SAM-3D, and demos include knocking over objects, breaking a mirror, and simulating an earthquake. Paper, code, and benchmark are released.

Original post →

More from Multimodal

Multimodal channel →