dots3-note's Test-Time Learning and TEMPO Recursive Self-Evaluation
dots3-note can self-correct and reuse experience in unfamiliar environments, even generalizing to untrained games. Its TEMPO mechanism chunks dozens-of-hours rollouts into macro-steps for process rewards, tackling sparse-reward long-horizon tasks.
2026-08-30 ~ 2026-08-30 · 2 related posts
- TEMPO: recursive self-critique splits day-long agent rollouts into macro-steps — SarahAnnabels · 2026-08-30
- dots3-note shows test-time learning: explores, self-corrects, reuses knowledge in unseen environments — SarahAnnabels · 2026-08-30