RT-SAFE benchmark: frontier VLMs hit 94.1% task success but only 0.7% finish safely
Lianhuiq · x · 2026-10-06
The SimWorld team introduced RT-SAFE, a benchmark evaluating embodied agents under real-time safety constraints while the world keeps moving, testing several frontier VLMs:
- Task success rose from 91.3% to 94.1%, but safe success (finishing without a safety event) collapsed from 19.8% to 0.7%, with 12.3× more collisions
- Agents can still complete tasks while becoming dramatically less safe
- More reasoning isn't the answer: longer inference itself creates physical risk when the world won't pause — latency is part of both intelligence and safety
More from Embodied
- Xitac tactile sensor uses magnetotactic gel over 3D hall array to capture finger forces — seanmcdonaldxyz · 2026-10-06
- ViDiHand: Video Diffusion Models Prove Surprisingly Good at 4D Hand Motion Reconstruction — ccloy · 2026-10-06
- RACE: 4x Longer Action Chunks for VLA Robots, 5x Less Idle Time — POSTECH · 2026-10-06
- InterMimicGen: Self-Evolving Motion Imitation Scales Humanoid Loco-Manipulation — Yucheng Zhang · 2026-10-06
- Humanoids casually roaming a San Francisco office ahead of speedrun Demo Day — tkexpress11 · 2026-10-06
- Hangzhou deployed Deep Robotics DR02 humanoids for rainy holiday traffic and tourist duty — CyberRobooo · 2026-10-06