NVIDIA's GATOR: generative 3D model + GPT-6 agent turns casual photos into simulation-ready objects

GATOR: Generative and Agentic 3D Object Reconstruction From Casual Images

Qirui Wu, Stan Birchfield, Hesam Rabeti, Angel X. Chang, Bowen Wen

cs.CV, cs.RO

2026-10-08

GATOR drafts posed, textured 3D objects from cluttered photos with a generative cascade, then lets a GPT-6 agent refine them, halving Pixal3D's chamfer distance on real tabletops.

What problem this solves

Turning casual photos into 3D assets that are actually usable: complete geometry, real textures, and the pose to place each object back into its scene. "Casual" is doing specific work here. No object isolation, no calibrated capture, no planned camera path. The target sits in tabletop clutter or a room corner, partly occluded, seen from a few irregular viewpoints, and cameras, depth, and masks all have to be estimated from the images.

The two existing paths each stall halfway. Generative reconstructors (LRM, TRELLIS.2, SAM 3D) fill unobserved backsides from learned priors, so a plausible-looking completion can diverge from the real object; when image, mask, and geometry arrive as independent conditioning streams, the denoiser has to infer how they correspond. Pure agentic pipelines, where GPT-6-Astra writes Blender code directly, control structure and materials but burn API time and cost building from scratch, and fine geometry suffers. In Figure 1 the agent misplaces a strainer's ear and simplifies its handle into crude shapes. GATOR chains the two: a generative model drafts a complete, scene-aligned asset, and an agent makes targeted repairs on top of it.

Method

The generative stage rides on TRELLIS.2's structure-geometry-appearance cascade:

The agentic stage hands the generated asset, the input images, and Blender to GPT-6-Astra (Codex CLI, extra-high reasoning effort). The agent decides what to fix: completing missing structure, repairing topology, correcting UVs and materials, including the legibility of printed text and logos. After each edit it renders the original and candidate under matched cameras and lighting and reverts anything that degrades agreement with the observations, keeping the best inspected asset within a fixed budget (ten minutes per object in the main experiments). The generative model is never retrained.

Results

BenchmarkMetricGATORBest baseline
Toys4K, 4 viewsCD / F10.009 / 0.757Pixal3D-MV 0.011 / 0.701
Tabletop (LM-O+HB+HANDAL)CD / F1 / LPIPS0.009 / 0.904 / 0.140Pixal3D CD 0.018, MV-SAM3D F1 0.780
ScanNet++, 229 objects, 49 scenesCD / F10.031 / 0.603RecGen 0.041 / 0.540
Pose, measured unaligned[email protected]96.36% HANDAL, 95.63% ScanNet++90.83% next best on ScanNet++

On ScanNet++, where cameras, depth, and masks are all estimated (SAM3 does the segmentation), GATOR cuts CD by 23% against runner-up RecGen and beats from-scratch GPT-6-Astra on every metric. The ablation adds one component at a time on 55 HANDAL objects: the modality mixer mostly drops ADD-SB from 0.057 to 0.037, which points to robustness against estimation error; text conditioning lifts all nine metrics (F1 0.872 to 0.930); agent refinement mainly helps appearance (PSNR 18.93 to 19.60).

The efficiency numbers are the most striking. Generation alone takes about 0.5 minutes (roughly 32 s of generative compute) and already outperforms GPT-6-Astra given a 20-minute budget. From scratch, GPT-6-Astra returns valid assets for 0 of 10 objects at one minute, 5 of 10 at two minutes, and all ten only from five minutes onward. More views keep helping: with two views, GATOR's CD is about half of Pixal3D's with eight.

Why it matters

For robot simulation, scene reconstruction, and digital-asset work, casual photos now yield simulation-ready objects with PBR materials and scene-relative pose, with no isolated or calibrated capture. Generalization to four real datasets happens with zero fine-tuning, resting on a synthetic-data-plus-error-injection recipe that is portable in itself.

Methodologically this is a clean case of "generative draft, agentic repair": the agent never models from scratch, it makes local edits anchored to instance-level geometry and pose, with time budgets, reverts, and no retraining of the generator. That division of labor should transfer to other 3D generation tasks. Honesty check: on clean synthetic data the margin over the strongest baseline is incremental (19.7% CD reduction on Toys4K). The real gaps open up in cluttered real scenes (tabletop CD halved) and in pose recovery, a dimension several baselines simply do not output.

Limitations

Terms

Source

What people are saying

Related papers

All paper explainers