A fine-tuned 9B beats a 31B model: 600 labels, $0.12, 91% accuracy

julsimon · x · 2026-10-08

Julien Simon fine-tuned Qwen3.5-9B on real banking support messages and scored it on 180 unseen messages. Base Qwen3.5-9B: 70.0%; Gemma 4 31B zero-shot: 77.8%; fine-tuned Qwen3.5-9B: 91.1%—for just $0.12.

How it was done:

His takeaway: when quality falls short, the reflex is a bigger model, but good data plus a small model wins more often—and runs faster and cheaper. The code is public so anyone can rerun it.

Original post →

More from coding & agent

coding & agent channel →