4B Model Hits 54% on IMO-ProofBench, Costing 3x Less Per Problem
SergioPaniego · x · 2026-08-22
A 4B parameter model achieved 54% accuracy on IMO-ProofBench, competing with models 7-50x its size while costing 3x less per problem.
Key techniques include:
- Pure post-training: No pre-training compute required.
- Distillation SFT: Supervised fine-tuning via knowledge distillation.
- GRPO: Reinforcement learning optimization with rubric rewards.
- Reasoning Cache: Test-time refinement using a caching mechanism.
More from Models
- Google Criticized: Gemini 3.7 Still Missing From Its Own Jules Agent a Week Later — brandon_galang · 2026-08-24
- Qwen 27B 3.8 low quantization tested: Q3 XXS works well locally — jeremyckahn · 2026-08-24
- Users notice significant quality shift in GPT-5.6 output — haider1 · 2026-08-24
- Ramp Stats: Anthropic Opus 4.8 and Sonnet 4.6 Lead Usage — vista8 · 2026-08-24
- Tencent Releases UI-Mate-27B, a Desktop GUI Agent Model — tencent · 2026-08-24
- ConvRot Quant joins llama-cpp: Q6 accuracy nears Q8 quality — giveen · 2026-08-24