Code4Scene benchmark: 190 Unreal Engine scenes test 14 coding agents on 3D work
kwangmoo_yi · x · 2026-10-02
Two related papers target coding agents for 3D scenes:
- Code4Scene (arXiv:2609.36777): a benchmark of 190 Unreal Engine human-assembled scenes evaluating two settings — construction from open-ended language specs, and editing that must recover a target scene from reference images. It scores the generated engine-native scene on task fulfillment, artifact integrity, and static physical validity rather than code or renders.
- Across 14 coding-agent configs on the 95-case public set, construction and editing are strongly correlated but not interchangeable (Spearman ρ=0.78): Claude Fable 5.1 leads construction, Gemini 3.8 Flash leads editing, GPT-6 Astra narrowly leads overall.
- A companion paper, LEGO-Anything, uses coding agents for 3D scene reconstruction. The author notes Gemini 3.8 was already strong before Astra.
Related event: Code4Scene Benchmark Tests Coding Agents on 3D Scene Building(3 posts)→
More from coding & agent
- Dev says Claude and Claude Code are getting worse day by day — Pavan_Belagatti · 2026-10-02
- Dev hooks Reels into Claude Code so he can doomscroll while the agent codes — yoimnotkesku · 2026-10-02
- Book a flight in 2 min or endure a 10-min AI agent call? Devs argue agent UX — jxnlco · 2026-10-02
- How llama.cpp Finds Safe Seams: Inside the Model-Loading Compatibility Hub — Mahmoud_Zalt · 2026-10-02
- Argo-Bench Pits Data Agents Against a 7.5B-Row Warehouse; Best Model Clears Only 34.8% of Tasks — textql · 2026-10-02
- Flexport launches MCP letting AI tools ship containers worldwide, customs included — ycombinator · 2026-10-02