Unsloth adds fine-tuning to turn Qwen, Llama and Gemma into small local decision models
blaizedsouza · x · 2026-10-09
- Unsloth now lets you fine-tune models like Qwen, Llama, and Gemma into "decision models" that make bounded decisions instead of generating text.
- A small decision head scores user-defined options and returns the selected answer with probabilities; training uses LoRA with a simple dataset format (input state, questions, gold answers).
- Models are calibrated on held-out data so probabilities better reflect actual correctness — useful for routing, classification, policy, and evaluation decisions with much smaller models.
- A companion article shows using such a decision model (Jev) as a review layer for a data analyst agent.
More from coding & agent
- Baton: open-source HTTP relay that lets any two AI agents talk, with signed transcripts — seanmcdonaldxyz · 2026-10-09
- Dev builds Baton, an open-source relay giving multiple AI agents shared chat rooms — seanmcdonaldxyz · 2026-10-09
- Meta paper introduces 'agent plasticity': measuring self-improving agent gains per dollar spent on learning — omarsar0 · 2026-10-09
- Claude drives AutoCAD via script, drafting a 524-element two-floor house in 88.6 seconds — natesiggard · 2026-10-09
- Google's Antigravity now supports Claude, billed by consumption via Gemini Enterprise — rseroter · 2026-10-09
- Open experiments: swapping the system prompt swings coding agent scores by 10% — yb2698 · 2026-10-09