Fine-tuning Qwen 3 4B on 100 zebra puzzles yields +31% on MATH-500
TGSCrust · reddit · 2026-09-11
A Hugging Face blog reproduces the pcss recipe: fine-tuning Qwen 3 4B Base on just 100 zebra logic puzzles yields a +31% improvement on MATH-500, with training taking only 6.5 minutes on a single H100/H200.
The standout is the extremely low barrier to entry: a full reproduction notebook is provided, letting anyone verify the result on one GPU and demonstrating how small, high-quality reasoning data can disproportionately boost small-model math ability.
More from Models
- V4.1 Ranks #5 on LiveBench, Tops Agentic Coding but Called Language-Skewed — teortaxesTex · 2026-09-11
- OpenAI ships GPT-Live prompting guide: copying your old prompts won't hit SOTA — craigsdennis · 2026-09-11
- GPT-6 Astra burns through user's weekly usage cap, forcing a wait until Monday — max_paperclips · 2026-09-11
- Wiz Launches Cyber Model Arena: Gemini 3.8 Flash Cyber Tops at 74.9% — rseroter · 2026-09-11
- Anthropic blocks minors from using Claude, HN debates age policy — petrusenko_max · 2026-09-11
- V4.1 hits #5 on LiveBench, tops Agentic Coding by 20 points over Astra — teortaxesTex · 2026-09-11