GAUGE: No Physics Engine Wins Across Contact, Cloth, and Soft Bodies

GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models

Shuai Wang, Yaxin Feng, Xuekun Jiang, Shihan Tian, Ningyu Yan, Xing Shen, Chaoyang Lyu, Hui Wang, Yunsong Zhou, Hanqing Wang, Jiangmiao Pang, Yang Xiang, Xing Gao, Chunhua Shen, Weinan Zhang

cs.AI, cs.CV, cs.RO

2026-08-06

GAUGE scores Isaac Sim, Genesis, and Newton on 22 real mocap families. No engine wins all regimes. World models can fit the right equation with the wrong acceleration or period.

What problem this solves

Embodied training already depends on simulation. Work such as SIMPLER showed that a well-aligned sim ranking can predict real-robot ranking. That only holds if motion, contact, and deformation match the world. Pretty pixels are not enough. An engine can look right while getting impulse, cloth flutter, and moduli wrong. Policies then exploit artifacts, and leaderboards mis-rank.

Video world models are being treated as implicit simulators: first frame plus a prompt, then a rollout. Most evaluations still ask whether the clip looks plausible. They do not say which law broke, or which parameter is off.

GAUGE uses one real experimental suite to diagnose both numerical engines and generative world models. Shanghai AI Lab leads, with HKUST, Shanghai Jiao Tong, and Zhejiang University. Xing Gao is corresponding author.

Method

The suite has 22 controlled task families: 8 rigid (collision, friction, momentum, oscillation), 1 rope, 6 textiles, 7 volumetric soft bodies. Materials include wood, plastic, metal, six fabrics, and soft/hard foam. Capture sits in a 2 m cube with 16 NOKOV Mars9H cameras at 180 Hz, millimeter-level. Rigid bodies get a 6-DoF frame from markers; cloth and foam get a tracked mesh. Each task has about 20 repeated trials, stored at 30 fps, with calibrated mass, friction, restitution, stretch and bend stiffness, Young's modulus, and Poisson's ratio. The authors put the full set at roughly 1,560 motion-capture trials.

The engine track scores Isaac Sim 6.0.0, Genesis 1.12.0, and Newton 1.3.0 on 14 families. Scenes are rebuilt from digital assets and mocap initial states. Physical parameters come from GAUGE calibrations; everything else stays default. Rigid bodies use PhysX, Genesis's native solver, and MuJoCo. Textiles use Surface FEM, PBD, and VBD. Soft bodies use FEM and explicit/implicit MPM. Error is computed on a generalized trajectory: centroid for rigids, Gaussian curvature at markers for cloth, triangle area for foam, then RMSE and DTW. Newton's cradle and the pendulum add longest stationary duration, momentum-transfer efficiency, period, and energy loss. Simulator error is divided by the real within-trial RMSE, so a value below 1 means the sim is tighter than real repeatability.

The world-model track stays on rigid tasks because current models still struggle with 3D consistency and deformation. Six models: Cosmos3-Nano, Cosmos3-Super-I2V, Wan-2.2, Wan-2.7, Seedance 2.0, Genie 3. Same frontal first frame, same prompt. SAM3 tracks the object; pixel centroids are scaled into a 2D world trajectory using known size. Scoring splits law form (R², quadratic-form improvement QFI) from parameter accuracy (acceleration, momentum transfer, period). Cosmos and Wan also run a paired ablation with a shared negative prompt.

Results

No engine wins the whole board.

Everyday contact and sliding are tolerable. Isaac Sim is lowest-error on slope contact, nonsmooth contact, and the turntable, with normalized RMSE / DTW of 0.17 / 0.61 on the turntable. Genesis wins the slope slider at 0.58 RMSE and 0.69 DTW. Impact and momentum open a much larger gap. The best bouncing-ball RMSE is still 15.63 times the real baseline. On Newton's cradle, Isaac and Newton post zero longest stationary duration and momentum-transfer efficiencies of 0.20 and 0.26; Genesis produces no valid rollout. Pendulum frequency looks easier: Isaac and Newton land at 1.10 and 1.09 times the real period, Genesis at 2.47. Over 3,000 frames, energy error sits between −0.041 and 0.034. Matching a low-frequency period does not mean impact or long-horizon energy is right.

SettingClosest engineHeadline numberAgainst
Turntable slidingIsaac SimRMSE 0.17Below real repeatability
Slope sliderGenesisRMSE 0.58Newton 1.95
Bouncing ballIsaac SimRMSE 15.63× baselineImpact still an order off
Cradle momentumNewtonMTE 0.26Real 1; Genesis no rollout
Textile flingGenesisRMSE 8.54Isaac Sim 128.26
Foam stretch/shear/twistGenesisStill 10× baselineNewton slightly better on bending

Quasi-static textile stretch can be brought near 1 for Isaac and Newton. Bending and flinging are not. Even the best volumetric numbers sit about an order of magnitude above real baselines.

For world models, correct shape with wrong scale is the main pattern. The best wood slope acceleration is 2.06 m/s² against a measured 2.58. Plastic and metal get as close as 0.75 and 0.43 versus 2.57 and 2.67. On the bouncing ball, Cosmos3-Super-I2V plus the negative prompt drops QFI to 12.50 while inferring 0.088 m/s²; gravity is 9.81, and Seedance's closest value is 1.84. Six of ten Newton's-cradle configurations never produce a usable sequence. The best momentum transfer is 0.76 from Wan-2.2 with the negative prompt. Wan-2.2 and Genie 3 fit the pendulum at R² 0.99 with periods 1.93 s and 1.90 s against 1.06 s. Wan-2.7's 1.83 s is the closest and still about 73% long. The negative prompt is not a reliable upgrade: it can cut bouncing-ball QFI and raise the wood slope QFI from 13.61 to 569.36.

Why it matters

This is a ruler that points at parameters, not vibes. Engine choice cannot be "does the cloth look soft." Isaac is stronger on several rigid contacts, Genesis on dynamic cloth and most soft bodies, Newton on selected deformation cases. If world-model papers only collect "looks physical" scores, they mix equation shape with physical scale. GAUGE splits those layers.

It is a diagnostic benchmark, not a new solver. Out-of-the-box defaults, no per-task tuning, will look harsh to teams that already calibrate digital twins. It is still closer to an actionable signal than a plausibility rating.

Limitations

The authors say the material set and parameter ranges are narrow, and that fluids and fluid-structure coupling are out of scope. Same-name fabrics still vary with finish, humidity, and wear. The world-model track uses 2D image trajectories, so distributed deformation, self-occlusion, and self-contact are not scored. Engines cover 14 of 22 families. Markers on light cloth are not strictly massless; the paper calls the effect negligible, but fast flings could still leak into the observation. Prompt-paired scores show that a single pretty clip is not a result.

Terms

Source

What people are saying

Related papers

All paper explainers