TEMPO evaluates progress in long-horizon tasks

thetripathi58 · x · 2026-08-26

Long-horizon tasks are hard because agents can spend hours acting without knowing if they are closer to the goal. TEMPO allows the model to switch between actor and critic at macro-steps to evaluate progress. For instance, in a knight-placement task, the Critic distinguished between trajectories with the same reward but different directions.

Original post →

More from coding & agent

coding & agent channel →