A fine-tuned 9B beats a 31B model: 600 labels, $0.12, 91% accuracy
julsimon · x · 2026-10-08
Julien Simon fine-tuned Qwen3.5-9B on real banking support messages and scored it on 180 unseen messages. Base Qwen3.5-9B: 70.0%; Gemma 4 31B zero-shot: 77.8%; fine-tuned Qwen3.5-9B: 91.1%—for just $0.12.
How it was done:
- Only 600 labeled examples, 3 epochs, 6 minutes of serverless training on Crusoe AI Intelligence Foundry, with the price quoted before the job starts
- The whole pipeline is one Python script written against the OpenAI fine-tuning API
- He also serves the adapter two ways—serverless vs. a dedicated NVIDIA H100—and compares them on camera
His takeaway: when quality falls short, the reflex is a bigger model, but good data plus a small model wins more often—and runs faster and cheaper. The code is public so anyone can rerun it.
More from coding & agent
- Managed Agents on Interactions API: one call spins up Antigravity agents in a remote sandbox — clmt · 2026-10-08
- Legacy billing rewrite estimated at 7-8 months shipped in six, 122 PRs in 90 days — alex_verem · 2026-10-08
- LangChain engineer built an ACP coding agent that replaced Claude Code for 9 months — Hacubu · 2026-10-08
- 30 Real Business Workflow Tests: Keep Agent Evaluation Simple — VibeMarketer_ · 2026-10-08
- Hybrid agent pattern: cloud Gemini plans, local Gemma swarm runs 97% of tokens offline — clmt · 2026-10-08
- Long-running agents suffer 'constraint amplification': a subtle form of context rot — generativist · 2026-10-08