KAIST Proposes SQAM: Scalar Adjoint Speeds Up Flow Policy RL, +18–35pp Success on OGBench
kaist-ai · hf · 2026-10-08
KAIST AI researchers introduce Q-learning with Scalar Adjoint Matching (SQAM), a method for off-policy RL fine-tuning of flow policies.
- Problem: Adjoint matching requires a vector-Jacobian product through the policy at every flow step, with cost growing with step count and policy size.
- Key finding: The batch-averaged velocity Jacobian of pretrained flow policies concentrates on its diagonal, enabling a closed-form scalar adjoint that scales the value gradient at the final action by flow time.
- Method: SQAM combines this scalar adjoint with a value penalty on policy-generated actions, which proves important for critic control.
- Results: On the four hardest OGBench domains, SQAM beats the strongest baseline per domain by 18 to 35 percentage points. It also fine-tunes a vision-language-action policy on a real bimanual robot, outperforming supervised fine-tuning on all three tasks.
More from Embodied
- Hedge fund analyst after robotics factory tour: 'It's not a decade away' — Rewkang · 2026-10-08
- ArUco-based visual tracking demo on an Air75 tiny whoop drone — k7agar · 2026-10-08
- Boston Dynamics names ex-Alexa chief Rohit Prasad as new CEO amid Atlas manufacturing push — lukas_m_ziegler · 2026-10-08
- Tesla patents tactile skin that conforms to curved robot hand surfaces, pointing at Optimus — CyberRobooo · 2026-10-08
- Paradromics BCI shows 4+ years of stable neural recording in sheep — KordingLab · 2026-10-08
- Wandercraft calvin-40 carries 40kg, spotlighting robots' gap on heavy 'dull, dirty, dangerous' jobs — chris_j_paxton · 2026-10-08