Code4Scene benchmark: coding agents still fail at building and editing Unreal Engine 3D scenes
Lianhuiq · x · 2026-10-01
Researchers introduce Code4Scene, a benchmark testing coding agents on constructing and editing Unreal Engine 3D scenes via code, evaluated across 14 agent configurations on 95 cases. Findings: spatial composition is the weakest construction skill for every agent, and 35.8% of edits that fully recover the target still introduce unintended changes elsewhere.
More from coding & agent
- Firecrawl's Alexandria hits #6 most popular ChatGPT plugin with 120+ data providers — devdigest · 2026-10-01
- Cloudflare launches agent-themed batch: pay-per-use gateway, AutoRouter, 6x faster containers — threepointone · 2026-10-01
- Blogger warns the cheap vibe-coding era is almost over — and the dots explain why — michalmalewicz · 2026-10-01
- Pocket FM's Memory System Matches Claude Code's 86.7% Accuracy at 1/21 the Cost — bigaiguy · 2026-10-01
- Ben Goertzel: AI agents are co-authoring formally verified software, not just patching bugs — bengoertzel · 2026-10-01
- Pedro Domingos: There are no software engineers anymore, we're all agent wranglers — pmddomingos · 2026-10-01