RT-SAFE benchmark: 94.1% of 8 frontier VLMs reach goals, only 0.7% finish with zero safety events
Lianhuiq · x · 2026-10-06
RT-SAFE is a new benchmark for evaluating embodied agents in a world that keeps moving while they reason. Built in Unreal Engine with a dynamic NYC scene, it asks VLM-powered agents to navigate sidewalks and crosswalks to a destination while recording 'passive collisions' during planning plus active collisions, hazard interactions, and traffic violations on execution.
- Scale: 5 city maps, 36 navigation routes, 16 available actions
- Evaluated: 8 frontier vision-language models
- Results: 94.1% reach the goal, but only 0.7% complete the route without any safety event
The gap highlights that current VLMs can navigate dynamic environments but rarely do so safely.
More from Embodied
- Stanford Open-Sources DITTO-X: Force-Feedback Teleop With Reverse Human Intervention — CyberRobooo · 2026-10-06
- $3,500 local AI PC under fire: only 24GB VRAM and priced below its own parts cost — BLUECOW009 · 2026-10-06
- Actual raw lidar image from Waymo's 6th-gen sensor suite shared online — reed · 2026-10-06
- Vibe Robotics: Artifact Arena makes frontier models engineer competing robots — kaixhin · 2026-10-06
- PointWAM: 3D world action model beats dexterous manipulation SOTA by 11.7 points — Chunghyun Park · 2026-10-06
- JEPA-TTT cuts latent world-model prediction error 83% under dynamics shifts — JHU · 2026-10-06