Robobench evals: Astra reasons with just 100 tokens vs 400-1k for rivals; Qwen 3.8 27B and GLM 5.3 Flash lead open models

arankomatsuzaki · x · 2026-09-28

Researcher arankomatsuzaki shares open-vs-closed model comparisons on Robobench and adjacent robotics tasks, measuring supervisor/executive capability. The standout finding beside the score gap: Astra consumes only 100 tokens for reasoning, while other models burn 400-1k even at low reasoning levels. Among open models, Qwen 3.8 27B and GLM 5.3 Flash performed best, close to Grok and Muse, while hy-embodied-vlm-1.0 and embodied r1.5 lagged far behind. He sees a narrow window to study the capability ramp of pre-physically-intelligent LLMs before the gap shrinks.

Related event: Researchers say now is the window for open models to catch up in robotics(2 posts)→

Original post →

More from Embodied

Embodied channel →