Ditto-Bench stress-tests GPT-6 Astra for robot control: simple tasks shine, physics stalls
anand_bhattad · x · 2026-10-09
JHU researchers Lin Long and Jaemin Cho built Ditto-Bench, where simple goals meet challenging physics (complex geometric constraints, precise contact, improvised tool use), to systematically test OpenAI's newly released GPT-6 Astra against MolmoAct2 for robot control.
Key ablations:
- Control interfaces: linear motion vs. action chunking
- Whether higher reasoning effort improves performance
- Removing memory (many VLAs have little or none)
- In-context learning from its own failures or others' demonstrations
Findings: Astra's semantic scene understanding and control logic are impressive on simple tasks — some have called it SOTA — but it gets stuck on challenging physics. Full code and blog post are public; a notable first independent benchmark of Astra's embodied capabilities.
Related event: Ditto-Bench Tests GPT-6 Astra's Robot Control Skills(4 posts)→
More from Embodied
- Open-Source Models Power Home Robot Tidying for Toddlers in Weeks, Not Years — chris_j_paxton · 2026-10-09
- Sony's quadruped walks on cracked pavement with passive-stability spherical feet — BLUECOW009 · 2026-10-09
- OpenAI's new Decisions API with image input tested on a MuJoCo robotics policy — RexDouglass · 2026-10-09
- TraceExtract open-sourced: data engine for µ0 world model trained on video with zero action labels — RexDouglass · 2026-10-09
- Robot fight tickets sell out in under 24 hours — cixliv · 2026-10-09
- NVIDIA shows frontier AI agents building Omniverse simulations from natural language — nordicinst · 2026-10-09