TTPO: Test-Time Policy Optimization Boosts Qwen3-1.7B from 38.0% to 45.2% Without Labels

JFPuget · x · 2026-08-29

TTPO (Test-Time Policy Optimization) lets LLMs keep improving at inference time without any ground-truth labels.

Related event: Alibaba and ZJU Propose TTPO for Label-Free Test-Time Self-Improvement in LLMs(3 posts)→

Original post →

More from Research

Research channel →