TEMPO: recursive self-critique splits day-long agent rollouts into macro-steps

SarahAnnabels · x · 2026-08-30

dots3-note addresses sparse rewards in long-horizon tasks with TEMPO: a rollout can take tens of hours, making the final reward too late to attribute credit. TEMPO breaks the trajectory into macro-steps; at each step the same model switches from actor to critic—reasoning over its state, calling tools, and estimating whether it is making real progress, forming recursive self-critique.

Related event: dots3-note's Test-Time Learning and TEMPO Recursive Self-Evaluation(2 posts)→

Original post →

More from Models

Models channel →