TMRL Improves Robot Policy Training

abhishekunique7 · x · 2026-07-10

A repost introduces TMRL: the author argues that RL cannot teach models to solve problems they have never encountered before, a limitation that also applies to robotic policies. This method injects diffusion noise into the policy during the pre-training phase to broaden the action distribution, thereby leaving room for exploration in subsequent RL. It also claims to enable VLA fine-tuning on a real robot within one hour.

Original post →

More from Embodied

Embodied channel →