2,000+ navigation tasks across 133 environments test how far LLMs are from zero-shot robot control

PaulYacoubian · x · 2026-09-29

Researchers gave Jev and other models a robot body and benchmarked them against Dimcode, Astra, Fable, Opus and 5.6 on 2,000+ navigation tasks across 133 real and simulated environments, covering navigation, spatial reasoning and world geometry.

Graded on speed, cost, tokens, collision count and path quality, the study asks how far language models are from zero-shotting complex real-time control tasks. The full dataset, code and paper are open source.

Original post →

More from Embodied

Embodied channel →