A 4B Model Trained for Under $500 Beats a 235B Sibling on Financial QA
AI Engineer · youtube · 2026-10-10
At AI Engineer World's Fair, Snorkel AI's Charles Dickens described how Snorkel and UC Berkeley's Sky Computing Lab trained a Qwen3 4B model with RL that scored 60% on financial QA versus 51% for its 235B sibling, for under $500 in training cost.
- Built the FinQA dataset from SEC 10-K filings with three layers of expert verification.
- Failure modes observed: hallucinated table schemas, context flooding, and repeated failed strategies — the real bottleneck was tool-use discipline, not reasoning depth.
- Trained with the open-source rLLM framework using a simple pass/fail reward; skills transferred to harder multi-table questions without hurting general tool use.
- Surprisingly, simpler training data and simpler binary rewards worked best. Model released as rLLM-FinQA-4B on Hugging Face.
More from coding & agent
- Alchemy + Cloudflare combo enables blazing-fast iteration, no local dev server needed — samgoodwin89 · 2026-10-10
- Cursor's Charlie Marsh reveals his AI code acceptance rate is 32.0% — keyanzhang · 2026-10-10
- Liquid AI's decision model d1 lands on Vercel AI Gateway with vision support — maximelabonne · 2026-10-10
- Voyager's Next Release Will Let You Build Playable Games on X — ravisparikh · 2026-10-10
- Basalt: open-source Blackwell inference engine hits 665 tok/s, 2.6x Strata on a 5090 rig — jesdga95 · 2026-10-10
- "Claude Code is the greatest video game of all time," quips developer — jarrodwatts · 2026-10-10