Coding Agents Can Author 3D Scenes but Not Control Them: Best Repair F1 0.527 on 190 UE Cases

Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes

Xiaokang Ye, Siddhant Hitesh Mantri, Zimeng Chen, Edward Zhang, Zhaoxu Zheng, Yuanheng Li, Yizhao Chen, Tianyang Huang, Lianhui Qin

cs.AI

2026-09-29

14 coding agents on 190 Unreal scenes: construction and editing correlate (ρ=0.78) but diverge; best edit Repair F1 is 0.527, and 35.8% of correct edits cause unintended changes.

What problem this solves

Coding agents can already write and execute the scripts that assemble a 3D scene in Unreal Engine. Whether the scene is actually right is a different question. A generated scene is a persistent, executable artifact, and a convincing render can hide incorrect spatial relations, objects intersecting each other, or modifications the request never asked for.

Standard evaluation habits miss this. Pipelines tend to score the generated code or a rendered view, and both skip the scene state actually sitting in the engine. Code4Scene targets that gap. It does not score the code or the renders; it inspects the engine-native scene.

Method

The benchmark has 190 Unreal Engine cases built from human-assembled scenes, run under a shared execution interface in two settings:

Scoring targets the engine-native scene along three axes: task fulfillment, artifact integrity (asset references and object structure intact), and static physical validity (the abstract names intersecting objects among the failures it catches). The load-bearing choice is where the measurement happens. Judging code style or render quality says nothing about spatial correctness; reading the object list and transforms straight from the engine says what is actually there.

Results

Four numbers carry the paper:

FindingValue
Spearman ρ between construction and editing ranks (14 configs)0.78
Best Repair F1 on editing0.527
Edits that fully recover the target yet change the scene elsewhere35.8%
Public set / total cases95 / 190

Three models split the crowns: Claude Fable 5.1 leads construction, Gemini 3.8 Flash leads editing, and GPT-6 Astra narrowly leads overall. Spatial Composition is the weakest construction category for every agent, which points to a shared weakness rather than one model's quirk.

A 0.78 rank correlation says the two skills travel together without being interchangeable. The editing ceiling is the sharper finding: the best configuration repairs barely half of what it should, and more than a third of the edits that do recover the target still touch something they were not supposed to. The abstract gives no absolute per-model scores, no metric formulas, and no list of the 14 configurations; those live in the full text, which was not retrieved for this write-up.

Why it matters

For anyone pointing agents at 3D content production, game-level assembly, simulation environments, training data for embodied AI, this is a capability map. The practical read: construction is close to usable, editing is not, and no single model dominates, so tooling choices should be made per task. The more durable idea is the evaluation entry point. Any pipeline that generates 3D content with an LLM can copy it: render quality does not certify spatial correctness, and checking engine-native scene state does.

Limitations

The paper's own conclusion is the cold shower: plausible 3D generation and reliable spatial reasoning and state control remain different things. From the abstract, validity checks are static only, and the scenes are human-assembled, so scale and diversity are capped by human effort. Open questions the abstract does not answer: whether failures from reading the reference images are separated from failures of spatial reasoning; whether results hold on the withheld half of the 190 cases; and how stable a rank correlation computed over 14 configurations really is. Confirming any of these requires the full text.

Terms

Source

What people are saying

Related papers

All paper explainers