β-OPSD lifts Qwen3-1.7B by 9.16 points on AIME 2024 avg@12
furongh · x · 2026-08-04
- On Qwen3-1.7B, β-OPSD improves AIME 2024 avg@12 from 44.2 with vanilla OPSD to 53.3, a 9.16-point gain.
- Across AIME 2024, AIME 2025, and HMMT 2025, the method gains +5.74 points on average over vanilla OPSD and beats GRPO on average.
- The slide argues that the method is more stable at every model scale shown: Qwen3-8B, Qwen3-4B, and Qwen3-1.7B all improve over vanilla OPSD.
Related event: β-OPSD Boosts Qwen3 Reasoning Performance, Ablation Reveals Key Mechanisms(2 posts)→
More from Research
- Experiment finds the method hurts in mature action-policy settings — YouJiacheng · 2026-08-04
- Roomer repairs 3D indoor layouts with object-grounded local edits — cn-scut · 2026-08-04
- ScrambleToolBench finds agents still brute-force tools after the map changes — declare-lab · 2026-08-04
- Hidden future trajectories make autonomous-driving VLMs reason more faithfully — Buaa1 · 2026-08-04
- WCM boosts VLA robot RL with a world-model critic and 149-task wins — OpenMOSS-Team · 2026-08-04
- StyleForge uses counterfactual reasoning to make room layouts more coherent — cn-scut · 2026-08-04