Fine-tuning Qwen 3 4B on 100 zebra puzzles yields +31% on MATH-500

TGSCrust · reddit · 2026-09-11

A Hugging Face blog reproduces the pcss recipe: fine-tuning Qwen 3 4B Base on just 100 zebra logic puzzles yields a +31% improvement on MATH-500, with training taking only 6.5 minutes on a single H100/H200.

The standout is the extremely low barrier to entry: a full reproduction notebook is provided, letting anyone verify the result on one GPU and demonstrating how small, high-quality reasoning data can disproportionately boost small-model math ability.

Original post →

More from Models

Models channel →