BuildingBench Tests Whether Coding Agents Can Turn Images Into 3D Buildings

Lianhuiq · x · 2026-09-17

BuildingBench is a new benchmark testing models' spatial understanding and reasoning alongside coding ability: agents write Python code to generate 3D meshes and textures from real-world images, no external 3D software needed. Early results: GPT-6 Astra scores highest at 80% lower cost than Fable 5.1, while DeepSeek V4.1 Flash delivers strong quality at a median $2.31 per building. Positioned as a step toward coding agents building reusable worlds for design, games, and simulation.

Related event: Enactra Launches BuildingBench for Code-Driven 3D Building Generation(4 posts)→

Original post →

More from Research

Research channel →