Label-free TTRL with ~100K bias params lifts Qwen2.5-7B to 76.67% on MATH-500
OLAResearchX · hf · 2026-10-01
Label-free bias-only TTRL
- Test-time RL without labels usually optimizes many parameters; this work shows adaptation emerges even with severely restricted reward and optimization space.
- Method: majority-vote pseudolabels as rewards, optimizing only 100K bias parameters with the backbone frozen.
- Results: Qwen2.5-7B reaches 76.67% on MATH-500, slightly beating a labeled bias-steering reproduction while optimizing 76,000x fewer parameters than full-parameter TTRL.
- The same procedure improves MathVista, AI2D, LogicVista and MMAU; steering vectors transfer to 4,500 held-out MATH problems.
- Analysis: majority-vote reliability improves with rollout consensus; bias subspaces with greater accessible gradient energy are more trainable.
More from Research
- How an internal 1945 EDVAC draft made von Neumann architecture famous — burny_tech · 2026-10-01
- Stanford's Kundaje calls out RNA model renaming: 'classical fine-tuning isn't a new model' — anshulkundaje · 2026-10-01
- SpikingBrain fuses linear attention with spiking neurons for zero-latency edge LLMs — gekobraa · 2026-10-01
- Ben Recht: utility maximization is inescapable—we must learn when it's misapplied — beenwrekt · 2026-10-01
- SCOPD self-distillation recovers 92% of full-context VLM accuracy with 90% fewer visual tokens — CSProfKGD · 2026-10-01
- Xiaomi MiMo-V2.6 report: one mixed GRPO run across all domains, $2.6M for Pro — SergioPaniego · 2026-10-01