VIGA agent rebuilds images into editable Blender scenes via multimodal inverse-graphics loop
Michael_J_Black · x · 2026-09-12
Researchers from UC Berkeley, CMU, and Max Planck present VIGA (Vision-as-Inverse-Graphics Agent) at ECCV: a multimodal agent that reconstructs input images as editable scene programs in Blender via an analysis-by-synthesis loop with interleaved multimodal reasoning and evolving contextual memory. It can build scenes from primitives or leverage tools like Meshy and SAM-3D, and demos include knocking over objects, breaking a mirror, and simulating an earthquake. Paper, code, and benchmark are released.
More from Multimodal
- Building cool 3D things that show up in your living room with Astra, Blender, Unity — Scobleizer · 2026-09-12
- MiniMax H3 video gen on rented GPUs: 4090 does a clip for $0.013, L40S matches 4090 speed — Worldly_North_7213 · 2026-09-12
- MiniMax H3 Director on a 16GB RTX 5060 Ti: 23-second video in 24 minutes — thatguyjames_uk · 2026-09-12
- Dev Builds Dream Game Solo: GPT-6 Auto-Rigs Creature Animation From Tripo Meshes — Vjeux · 2026-09-12
- Dev builds game world you can reshape by voice in real time using OpenAI's new GPT-Live-1 API — TheMoonMidas · 2026-09-12
- First Impressions of YuE2 Music Generation in ComfyUI — Lividmusic1 · 2026-09-12