TEMPO Tops ARC-AGI-3 by Switching One Model Between Actor and Critic
TEMPO tops ARC-AGI-3, gaining 31.5% over the base checkpoint and 20.6% over GRPO, by splitting long trajectories into macro-steps and switching a single model between actor and critic roles, with critics proving more informative than environment rewards.
2026-08-20 ~ 2026-08-20 · 3 related posts
- TEMPO switches the same model between actor and critic to value long-horizon rollouts — omarsar0 · 2026-08-20
- TEMPO Outperforms Baseline by 31.5% on ARC-AGI-3 — omarsar0 · 2026-08-20
- Critic vs. Environment Rewards: Hidden-Rule Game Case — omarsar0 · 2026-08-20