Stanford's three-tier controller gets GPT-6 Astra to 79.13% on RoboMME with just 3.63 calls per episode
DJiafei · x · 2026-09-22
A Stanford team (Bingao Chen, Haoquan Fang, Prof. C. Karen Liu) presents a memory-augmented manipulation system built on GPT-6 Astra, extending their earlier SAM2Act work on spatial memory for robots.
- Three-tier architecture: System 1 is a low-level vision-language-action (VLA) model executing actions; System 2 is GPT-6 Astra doing high-level planning and grounded subtask generation; System 1.5 is a small learned vision-language model monitoring subtask completion.
- Cost-saving key: the small model decides when to re-invoke Astra, separating frequent progress checks from expensive high-level reasoning instead of calling Astra at every action chunk.
- Results: 79.13% success across 800 official RoboMME test episodes in 16 tasks, averaging only 3.63 Astra calls and 80.84 seconds per episode—nearly matching the 84.08% oracle-guided upper bound (GroundSG+Oracle).
- Code is open-sourced.
More from Embodied
- Cybercab #67 joins Tesla's Austin robotaxi fleet — EricETesla · 2026-09-22
- Robot beats up human in a cage match in San Francisco — Charuru · 2026-09-22
- Decentralized multi-humanoid transport: one policy, no comms, pinch-lift-move (decMHT) — tweetsatpreet · 2026-09-22
- Saudi brand Ceer unveils Exobot, an AI-designed concept humanoid — aziz4ai · 2026-09-22
- Weekly AI Recap: Helix 2.5 Lifts Household Robot Success to 56%, World Models Meet Robot Control — TheTuringPost · 2026-09-22
- SemiAnalysis tears down A20 on TSMC N2, DRAM revealed in iPhone 18 Pro Max package — dylan522p · 2026-09-22