GPT-6 Astra Autonomously Trains a Pen-Spinning RL Policy in 36 Hours

JasonMa2020 · x · 2026-09-17

Walter Zhu's fourth GPT-6 Astra test: a single prompt asking for pen spinning with a dexterous Sharpa hand, RL training in Isaac Lab, a self-created pen mesh, and a visualization video — with freedom to search the web and download papers. After running autonomously for a day and a half (including policy training), it produced a working pen-spinning RL policy and demo video, prompting Jason Ma to note Eureka was ahead of its time.

Original post →

More from Embodied

Embodied channel →