OPSA boosts AIME24 by 35 points using self-entropy without teacher distillation
heghbalz · x · 2026-09-01
New research questions on-policy distillation, suggesting gains come from suppressing unlikely tokens rather than teacher knowledge. The authors introduce OPSA (On-Policy Self-Adaptation), a teacher-free method that assigns entropy-adaptive negative advantages to low-probability tokens. This self-supervised approach achieved a +35 point improvement on the AIME24 benchmark.
More from Models
- Qwen 3.8 27b oneshots a Super Mario clone in single attempt — zannix · 2026-09-01
- Company offers unlimited Claude, yet most employees stick to default Sonnet 5 — LegitimateLength1916 · 2026-09-01
- Local LLM SVG Generation Test: Qwen 3.8 Leads — pmigdal · 2026-09-01
- Mystery Gemma Model Spotted on Arena Leaderboard — Hot_Example_4456 · 2026-09-01
- antirez shows DeepSeek v4 Flash vision running fast locally on an M5 Max; Metal/CUDA/ROCm support nearly ready — antirez · 2026-09-01
- Achieving 44% on ARC-AGI-1 Benchmark for 67 Cents — porridgeraisin · 2026-09-01