BuildingBench tests coding agents on 3D building generation, with up to 80% cost gaps
ZhitingHu · x · 2026-09-17
EnactraAI introduced BuildingBench, a benchmark testing whether coding agents can turn real-world images into coherent 3D buildings, aiming at reusable worlds for design, games and simulation. GPT-6 Astra scores highest at 80% lower cost than Fable 5.1; DeepSeek V4.1 Flash delivers strong quality at a median $2.31 per building. Interactive leaderboard and GitHub are live.
More from coding & agent
- YC built AI versions of its partners on GLM-5.2, cutting latency 31% vs OpenAI — ycombinator · 2026-09-17
- Dev ports bash-tool project to Claude Managed Agents, sandbox spins up independently — trq212 · 2026-09-17
- Developer Reverses Course: MCP Is Now Better Than CLIs for Most Integrations — trq212 · 2026-09-17
- Shape Claude's Tools the Way You Want Instead of Hiding Behind Indirection — trq212 · 2026-09-17
- Out of Codex quota? Signing in with a new Plus account preserves all your chats — TheMoonMidas · 2026-09-17
- Vercel's fx may switch safety reviewer to Jev: 5-18x faster than GPT Luna — thesaraharminta · 2026-09-17