2,000+ navigation tasks across 133 environments test how far LLMs are from zero-shot robot control
PaulYacoubian · x · 2026-09-29
Researchers gave Jev and other models a robot body and benchmarked them against Dimcode, Astra, Fable, Opus and 5.6 on 2,000+ navigation tasks across 133 real and simulated environments, covering navigation, spatial reasoning and world geometry.
Graded on speed, cost, tokens, collision count and path quality, the study asks how far language models are from zero-shotting complex real-time control tasks. The full dataset, code and paper are open source.
More from Embodied
- Robotics startup Roboteur hiring gripper/hand specialist, more home demos next — neurosp1ke · 2026-09-29
- Tensr builds robot factories that make robots, already supplying humanoids and ISS hardware — ycombinator · 2026-09-29
- ProHand Gen 2 robot hand launches at IROS 2026 with 2M-cycle durability and +66% abduction — chris_j_paxton · 2026-09-29
- AMD Acquires Fei-Fei Li's Spatial Intelligence Startup World Labs — amrrs · 2026-09-29
- Figure CEO Brett Adcock says goodbye to the F.02 humanoid robot — adcock_brett · 2026-09-29
- Self-supervised multisensory pretraining robot RL paper wins best student paper at IROS — GeorgiaChal · 2026-09-29