BuildingBench tests coding agents on 3D building generation, with up to 80% cost gaps

ZhitingHu · x · 2026-09-17

EnactraAI introduced BuildingBench, a benchmark testing whether coding agents can turn real-world images into coherent 3D buildings, aiming at reusable worlds for design, games and simulation. GPT-6 Astra scores highest at 80% lower cost than Fable 5.1; DeepSeek V4.1 Flash delivers strong quality at a median $2.31 per building. Interactive leaderboard and GitHub are live.

Related event: Enactra Releases BuildingBench for Evaluating Coding Agents on 3D Building Tasks(2 posts)→

Original post →

More from coding & agent

coding & agent channel →