TTPO: Test-Time Policy Optimization for Label-Free Math Reasoning

Aozhe Wang · hf · 2026-08-28

TTPO enables label-free test-time training for mathematical reasoning by asymmetrically distilling agreeing rollouts and penalizing disagreeing ones. It matches supervised performance on math tasks.

Original post →

More from Research

Research channel →