BuildingBench Tests Whether Coding Agents Can Turn Images Into 3D Buildings
Lianhuiq · x · 2026-09-17
BuildingBench is a new benchmark testing models' spatial understanding and reasoning alongside coding ability: agents write Python code to generate 3D meshes and textures from real-world images, no external 3D software needed. Early results: GPT-6 Astra scores highest at 80% lower cost than Fable 5.1, while DeepSeek V4.1 Flash delivers strong quality at a median $2.31 per building. Positioned as a step toward coding agents building reusable worlds for design, games, and simulation.
Related event: Enactra Launches BuildingBench for Code-Driven 3D Building Generation(4 posts)→
More from Research
- Harvard study suggests smartphones could predict suicide risk with striking accuracy — pshrink · 2026-09-17
- Study of 7 models across Claude Code, Codex, Pi: harness barely affects success but swings cost — DavideCrapis · 2026-09-17
- Neuroevolved value function solves Tetris at world-record speed in your browser — NathanWilbanks_ · 2026-09-17
- Polymarket atomicity gap: 1.8M reverted trades expose ghost-filled order attack — chaumian · 2026-09-17
- NeurIPS 2026 opens financial aid and volunteer applications, due Oct 6 — NeurIPSConf · 2026-09-17
- Classifying 1,018 AI Papers for $4: A Two-Model Pipeline at 256ms Median Latency — nutlope · 2026-09-17