TTPO paper: Qwen3-1.7B lifts accuracy from 38.0% to 45.2% without ground-truth labels
TheTuringPost · x · 2026-09-02
Turing Post highlights TTPO (Test-Time Policy Optimization), a paper showing models can learn at test time without any ground-truth answers.
How it works:
- The model never sees true answers — it uses a majority vote across its own attempts as pseudo-labels
- Solutions agreeing with the majority are treated as likely useful; disagreeing ones as likely wrong — but TTPO penalizes only the most suspicious tokens, not whole solutions
- This asymmetric approach reduces the risk of reinforcing wrong answers versus treating the majority as definitely correct
Results:
- Qwen3-1.7B's average accuracy improved from 38.0% to 45.2% with no answer labels
- Gains transferred across math benchmarks, suggesting broader learning beyond one dataset
Why it matters: models can keep learning even where no answer key exists.
Related event: TTPO Enables Test-Time Learning Without Ground Truth(2 posts)→
More from Models
- AxiomProver tops LeanEval, the last unsaturated math formalization benchmark — BenBlaiszik · 2026-09-03
- Muse Spark 1.3 calls user 'Judah' then denies it, users report odd behavior — fragment_me · 2026-09-03
- Google AI Mode shows zero citations on high-level TOFU queries, SEO tests find — gaganghotra_ · 2026-09-03
- Fable 5.1 halves agent failure rate to 7% with 0.7% hallucinations, at 1.8x the cost — ryanshrout · 2026-09-03
- Seroter Daily #859: Gemini 3.8 Flash, agent telemetry, and 7 agent skill patterns — rseroter · 2026-09-03
- Insider leak: OpenAI's Astra tested as 'ultima-alpha' and 'vega-alpha' checkpoints — Ok_Display_3159 · 2026-09-03