Robobench evals: Astra reasons with just 100 tokens vs 400-1k for rivals; Qwen 3.8 27B and GLM 5.3 Flash lead open models
arankomatsuzaki · x · 2026-09-28
Researcher arankomatsuzaki shares open-vs-closed model comparisons on Robobench and adjacent robotics tasks, measuring supervisor/executive capability. The standout finding beside the score gap: Astra consumes only 100 tokens for reasoning, while other models burn 400-1k even at low reasoning levels. Among open models, Qwen 3.8 27B and GLM 5.3 Flash performed best, close to Grok and Muse, while hy-embodied-vlm-1.0 and embodied r1.5 lagged far behind. He sees a narrow window to study the capability ramp of pre-physically-intelligent LLMs before the gap shrinks.
Related event: Researchers say now is the window for open models to catch up in robotics(2 posts)→
More from Embodied
- Stanford runs OpenAI's GPT-6 Astra on a Unitree G1 robot to tidy an unseen kitchen — alex_verem · 2026-09-28
- ETH Zurich trains robotic hand to walk on its five fingers across 14 surfaces — burny_tech · 2026-09-28
- Builder brings a physical Codex Pet to San Francisco for OpenAI Dev Day — paw_lean · 2026-09-28
- Researchers debate GPT-6 Astra: a generalist VLA that could beat custom VLAs — YouJiacheng · 2026-09-28
- Eric Jang: now is the window to test OSS models on robotics before the gap closes — ericjang11 · 2026-09-28
- 10 NeurIPS 2026 robotics papers to watch: VLA inference and long-horizon tasks — shaohua0116 · 2026-09-28