SimWorld's RT-SAFE benchmark: frontier VLMs' safe success plunges to 0.7% in real time
Lianhuiq · x · 2026-10-06
SimWorld introduces RT-SAFE, a benchmark evaluating embodied agents under real-time safety constraints. Testing frontier VLMs showed task success at 94.1% but safe success collapsing from 19.8% to 0.7%, with 12.3× more collisions. Key insight: longer reasoning itself creates physical risk when the world keeps moving—latency is part of both intelligence and safety.
More from Embodied
- The robotics data business's biggest bottleneck is operational, not technical — paigeinsf · 2026-10-06
- GPT-2-powered robot Samson explores, gets a little confused but has fun — MikePFrank · 2026-10-06
- VC: almost none of real robot deployment data flows back into training — carrycooldude · 2026-10-06
- RealtimeWAM: One-Step Asynchronous World Action Model Delivers 25x Speedup with <1% Accuracy Loss — NanyangTechnologicalUniversity · 2026-10-06
- ACG-Bench Probes Dual-Arm VLA Generalization; AE-VLA Lifts Success from 3% to 21.5% — Zaibin Zhang · 2026-10-06
- ReSteer open-sourced: fixing VLA policies that ignore mid-execution instruction switches — siddkaramcheti · 2026-10-06