TTPO: Test-Time Policy Optimization for Label-Free Math Reasoning
Aozhe Wang · hf · 2026-08-28
TTPO enables label-free test-time training for mathematical reasoning by asymmetrically distilling agreeing rollouts and penalizing disagreeing ones. It matches supervised performance on math tasks.
More from Research
- Survey on LLM Agent Evaluation: Taxonomy and Enterprise Challenges — kalyan_kpl · 2026-08-28
- Building a Sci-Fi Movie RAG Agent in ~60 Lines of TypeScript — mastra_ai · 2026-08-28
- 2026 AI System Danus Reproduces Complex Matroid Theory Proof — maier_ak · 2026-08-28
- Reproducing Recirculation: Gemma 3 PPL Drops 23% Without Fine-tuning — CatAstro_Piyush · 2026-08-28
- Researchers suggest AI conferences should run at a loss to subsidize youth — 3scorciav · 2026-08-28
- Nanjing Univ Proposes HCL: Continual Learning for Agents Without Model Tuning — 机器之心 · 2026-08-28