LAC paper: shifting RL architecture burden to the critic cuts robot inference latency 4x
heghbalz · x · 2026-09-05
A new offline RL paper questions why actors bloat with slow diffusion/flow-matching models while the critic is discarded after training.
- In real-world robotics, multiple denoising steps per decision destroy control frequency
- LAC shifts the burden to training: a residual MLP backbone, n-step bootstrap targets, and categorical cross-entropy stabilize a deep critic against optimization collapse and bootstrap drift
- With the deep critic doing the heavy lifting, a bare-bones deterministic actor matches heavy generative baselines on OGBench
- Single forward pass yields up to 4x lower inference latency
More from Embodied
- EU AI Act high-risk rules for humanoid robots now live, mandate audit logs — sierracatalina · 2026-09-05
- whurley rides Tesla CyberCab, declares it will put Uber out of business — whurley · 2026-09-05
- Semiconductor insiders at SEMICON are blunt: "humanoids are stupid" — bookwormengr · 2026-09-05
- How to train a fighting robot: RL goes into the ring — cixliv · 2026-09-05
- Tesla: Cybercab trained on 16,000+ lifetimes of driving data — yunta_tsai · 2026-09-05
- Figure CEO Brett Adcock: perfection comes in F.04, not v1 — adcock_brett · 2026-09-05