Unsloth shows fine-tuning Qwen into a decision model in ~2 min on one DGX Spark, 81% accuracy
danielhanchen · x · 2026-10-10
- Unsloth released docs for training your own decision models and launched Unsloth Desktop, a local app to run and fine-tune models. Fine-tuned LLMs (Qwen, Gemma, Llama) score input options and return a calibrated choice instead of generating text.
- Using a Clef head with LoRA (r=64) for just one epoch lifts accuracy from 30-37% (chance level) to 78-81%; BANKING77 jumps from 1-7% to 58-74%.
- Costs are tiny: Qwen3.5-0.8B needs 4GB VRAM and 42 min; Qwen3.5-2B 8GB/40 min; Llama 3.2 3B hits 79% in 30 min. A demo trained on a single DGX Spark in 2 minutes.
- Training mixes 12 datasets plus typed-decisions; test sets (3,000 rows) were decontaminated against training data.
More from coding & agent
- WareTwin: Open-Source 3D Digital Twin Simulating 20 Warehouse AMRs on GitHub — rsasaki0109 · 2026-10-10
- WareTwin: Open-Source Real-Time 3D Digital Twin Simulating 20 Warehouse AMRs — rsasaki0109 · 2026-10-10
- Google adds "Code Comprehension" interview round: debug a real codebase with Gemini, and catch when it's wrong — burny_tech · 2026-10-10
- Our AI bill hit $11,400 with no attribution: a cautionary tale of unbounded retries and prompt bloat — vigilAPI · 2026-10-10
- Doubao canvas upgrade: one image grows into a 24,000px VI manual, e-commerce pages and decks — xiaohu · 2026-10-10
- Dev uses Codex to migrate a 5.5TB S3 setup to AWS — CtrlAltDwayne · 2026-10-10