Open Tool Reveals Pixel Metrics Fail to Rank World Models on Real Robot Video
georgia_bucea · reddit · 2026-08-14
The author developed an open-source tool called Worldproof to diagnose where world models (models predicting future frames) break down, leading to a counterintuitive and crucial finding.
Core Finding: Pixel Metrics Lack Discriminative Power
- On real SO-101 robot arm recordings, standard pixel metrics like SSIM and PSNR cannot effectively rank models.
- This happens because prediction errors do not grow significantly with the horizon. If predicting 6 steps ahead is no harder than 1 step, all models appear similarly mediocre.
The Usable Evaluation Window
- The author measured 48 steps on the DROID dataset and identified three regimes:
- Steps 1-3: Near-perfect ties across all models.
- Steps 4-24: Steep monotonic decline; the only stretch where models are actually separable.
- Step 28+: Prediction fully decorrelated, and all models tie again at the bottom.
Methodology & Advice
- When evaluating world models, researchers must measure their own usable window based on frame rate and task speed, rather than blindly inheriting default horizons from papers.
More from Embodied
- Elon Musk Hints at a 'Catgirl' Version of the Optimus Robot — justalexoki · 2026-08-14
- Humanoid Cleaning Service Launches in SF at $30/Hour per Robot — chris_j_paxton · 2026-08-14
- Teleoperating Robots is Like a Video Game: Gamers Learn Faster — Ronangmi · 2026-08-14
- Matic Robot Tests New Cue Feature — chris_j_paxton · 2026-08-14
- Open-Source Fluidd: Manage Multiple 3D Printers with Klipper Web UI — tom_doerr · 2026-08-14
- When Robots Build Robots, Manufacturing Stops Scaling Like Humans Do — TansuYegen · 2026-08-14