USTC's VLA-Precision: online RL for VLAs hits 98.3% success on precision chemistry tasks
ustc · hf · 2026-09-28
USTC researchers present VLA-Precision, an efficient real-world online RL framework for vision-language-action models, targeting two bottlenecks: unreliable value signals causing policy drift, and large-VLA overhead limiting throughput.
- ACoB algorithm: asymmetric co-bootstrapping across timescales — early intervention-guided behavioral learning rapidly improves policy and experience quality; global return propagation plus local preference ranking progressively calibrate value estimates while suppressing drift
- ACoB-Stream architecture: invariant-state decoupling and on-demand streaming deliver up to 10.9x throughput/compute efficiency gains
- Results: across 9 high-precision chemistry tasks and 4 robot embodiments, 98.3% mean success rate at 45.8 min/task, 27.6s episodes running 1.2x and 1.8x faster than VLA and RL baselines
Resources at vla-precision.github.io.
More from Embodied
- HSImul3R (ECCV 2026): simulation-ready human-scene reconstruction judged by physics, not looks — jiqizhixin · 2026-09-28
- World's largest humanoid robot livestream performance draws crowd of 10,000+ — kernelangus420 · 2026-09-28
- One-arm VR intervention during bimanual DAgger feels like an AI brain chip — neurosp1ke · 2026-09-28
- Robot training fix: cutting noise injection from 2.5% to 0.25% of joint range removes jerkiness — DominiqueCAPaul · 2026-09-28
- AHa-3D turns ordinary indoor videos into editable Blender 3D scenes — ducha_aiki · 2026-09-28
- First real-world RL fine-tuning on a VLA: policy turns jerky after ~10 rollouts — DominiqueCAPaul · 2026-09-28