USDCraft turns partial meshes into sim-ready joints, with at most a 10-point real-robot drop

USDCraft: Geometrically Grounded Programmatic Modeling of Articulated 3D Assets for Simulation

Chuanrui Zhang, Zaijia Yang, Duomin Wang, Lu Shi, Daquan Zhou, Ruihua Zhang, Ziwei Wang

cs.RO

2026-10-08

A pretrained LLM writes articulated USD from partial meshes, with no task-specific training. Part F1 is 84.8% on USDCraft-bench, and real-robot success drops by at most 10 points.

What problem this solves

A manipulation policy trained in simulation transfers only when the asset matches the real object's size, puts the joint where the hardware is, and carries mass, friction, and limits the simulator can run. Particulate, SIMART, and ArtLLM partition a static mesh and estimate joint axes. Holes and fused parts break that partition, and openings the scan never saw stay missing. Those models are trained on labeled assets, so uncommon mechanisms travel badly. Articraft and Procedura instead have an LLM write a program from scratch. Category coverage is wide, but nothing ties the dimensions to one physical instance, and the grasp point learned in simulation misses the real object.

Method

USDCraft uses a pretrained LLM with no task-specific finetuning. The model reads a source mesh and a reference image, picks its own tools, writes an executable asset program, compiles it, and revises it. Generation is the same loop without a mesh, conditioned on text or an image.

An LLM cannot read a mesh. Renders throw away metric scale, and a vertex list does not fit in context. Source geometry analysis samples the surface onto one metric grid, stacks slices along the vertical axis, and records bounds, protrusions, openings, and interior structure. Cells the source surface never touches are marked unknown, neither empty nor solid. A hole in a scan can be missing data, and the inside of a closed shell can be empty of material, so both stay open for the image and the object's function to fill.

Parts, joints, and physical quantities are parameters in the program, so an error is a local edit. Mechanical parts use CadQuery, organic shapes use a signed distance field, and textures live in the same code. Mass, friction, limits, and drives are written as PhysX parameters. Compilation adds collision geometry and exports USD that loads in Isaac Sim unchanged.

Geometric rechecking encodes each candidate on that same grid. Source cells the candidate misses are edits to make. Extra cells where the source is unknown may be inner walls or unscanned parts, so they are not all treated as errors. First-surface depth from six axis-aligned directions splits the mismatch: the median offset b is a front-back shift, and eprofile, the 90th percentile of the residual after removing that shift, is shape error. On the handle in the overview, b is 0.64 mm and eprofile is 3.41 mm, so the profile is what changes. Visual feedback then renders per-part colors at chosen joint poses and looks for missing parts, bad assemblies, and collisions that show up only in motion. With no mesh, generation depends on that check.

Results

USDCraft-bench has 60 agent-authored assets, 20 each from USDCraft, Articraft, and Procedura, plus 40 generated meshes and real scans. Inputs carry no part labels. Sol is GPT-5.6 at high reasoning effort. Astra is GPT-6 at low effort, and its scores are means of three runs. ArtLLM is scored only where P3-SAM segmentation succeeded.

MethodPart F1 (%)Static gIoUArticulated gIoUAxis error (°)Axis location error
ArtLLM (successes only)19.30.1940.18961.20.366
SIMART36.00.2820.27346.70.262
Particulate33.40.3000.28948.30.244
Articraft-Astra64.00.2040.19615.10.092
USDCraft-Sol79.80.5860.5689.20.068
USDCraft-Astra84.80.6940.6706.80.044

Both USDCraft settings lead on all nine metrics, and lower-effort Astra still beats higher-effort Sol. Segmentation cannot add parts the mesh never contained. Image-only Articraft reaches a static gIoU of 0.204. Part Chamfer distance drops from Particulate's 0.127 to Astra's 0.052.

Lightwheel has 243 human-modeled assets in 14 categories. With no extra prompt, Astra's part F1 is 87.0%, above Instruct-Particulate on the original meshes (83.4%) and Articraft (67.0%). Static gIoU is 0.590, fully articulated gIoU is 0.555, and joint location error is 0.009. Segmentation keeps a lower whole-object Chamfer, and Instruct-Particulate a higher mIoU, because those methods reuse the input surface while USDCraft rebuilds every part. Articraft's axis error is 7.7° against USDCraft's 10.2°, but movable-part recall is 66.8% versus 85.0%, and joint errors average only over matched parts. One shared note of articulation conventions for all 14 categories, with no instance labels, lifts F1 to 89.8% and cuts axis error to 4.5°.

With the backbone fixed, the geometry tools beat a general-purpose Mini Workflow on all nine metrics, most clearly on the joint axes. USDCraft-Sol beats Mini Workflow-Astra on every metric. Claude Opus 5 and Gemini 3.8 Flash, on the same harness, also beat Mini Workflow on most metrics. The extra analysis costs about one to two minutes on the GPT backbones. Astra averages 4.2 minutes and $0.85 per asset. Claude averages 27.5 minutes and $4.51.

An iPhone Pro scan supplied a drawer, a toaster, and a support box, about two minutes each. Every asset that could support the interaction contributed 200 simulated demonstrations. Diffusion Policy trained for 80 epochs at batch size 128 and ran on a UR7e with a Robotiq 2F-85, with no real-world finetuning. Each cell has 20 simulated trials and 20 real trials. Success means opening the drawer by at least 5 cm, pressing the lever by at least 4 cm, or turning the switch by more than 5°.

Asset sourceOpen drawer sim/realPress lever sim/realSwitch toaster sim/real
Articraft-Sol90% / 20%90% / 10%80% / 20%
Particulate0% / 0%50% / 10%0% / 0%
Mini Workflow-Astra100% / 0%90% / 60%85% / 70%
USDCraft-Astra100% / 90%90% / 85%85% / 85%

Sol and Astra lose at most 10 points. Articraft loses 60 to 80. Mini Workflow-Astra scores 0% on the real drawer: the handle is misplaced, contact is unsafe, and the protocol counts that as failure. Particulate recovers only the lever, and its bodies were fixed in simulation to reach the 50% simulated success.

On 50 text prompts and 50 image prompts, an anonymous Sol/high judge scored five axes from 1 to 5. USDCraft holds the top score on every axis and the shortest runtime. Astra's median is 4.9 minutes. Procedura-Sol's is 58.7. The same toolkit produced USDCraft-10k, 10,000 assets across more than 500 categories, 70% from text and 30% from images.

The authoring toolkit alone drops axis error from 21.1° to 9.2°. Removing source geometry analysis drops F1 from 84.8% to 82.0% and static gIoU from 0.694 to 0.616, and raises axis error to 10.2°. Removing visual feedback leaves gIoU at 0.694 and drops F1 to 83.6%. An extra image-only observer lowers all nine metrics.

Why it matters

The work was done in an NVIDIA internship. The USD already carries collision geometry and PhysX parameters, so it loads into Isaac Sim without hand cleanup. The real-robot drop stays within 10 points because the handle, lever, and switch follow the measured mesh. A new category does not need a newly trained articulation network. The library already holds 10,000 assets.

The loop still runs in minutes. There is no timing comparison against a feed-forward segmenter such as Particulate. A complete mesh in a familiar category is still a segmentation job. A broken scan that has to train a policy is where revising a measured program fits.

Limitations

Fine lattices do not come back faithfully. Racket strings and shopping-cart wires keep a rough outline, but spacing and local connections miss the source. Deformable parts are stored as rest geometry. Revolute joints on umbrella ribs do not model canopy tension, and flexible-body dynamics were not validated in Isaac Sim.

Twenty of the 60 agent-authored assets were generated by USDCraft itself. After those references are removed, Astra's F1 is still 83.9%. After every agent-authored reference is removed, it is 89.8%, still first on all nine metrics. ArtLLM's 19.3% excludes 28 segmentation failures. Generation scores come from Sol/high, the same model family as USDCraft-Sol, and the judge sees kinematic snapshots rather than a physics rollout. The robot study uses three objects, three tasks, and 20 real trials each, and every method's supports are USDCraft-Astra reconstructions. Friction and mass come from authoring guidelines, not from measurement. The main table normalizes each object on its own. In the reference frame, Astra's F1 remains 83.3% and still leads all nine metrics, while image-only Articraft falls from 64.0% to 37.4%.

Terms

Source

Related papers

All paper explainers