Confidence calibration beats raw accuracy: Nimble 9B automates 2.7x more routing traffic than Qwen 3.5-9B
AgitatedUsual8995 · reddit · 2026-10-09
A Reddit user benchmarked Nimble 9B against its base model Qwen 3.5-9B on 900 real consumer complaints routed to 9 teams.
- Overall accuracy is nearly identical: 77.1% vs 76.3%
- Latency is the same: 2.25s vs 2.07s on the same hardware (8-bit)
- The big gap is confidence calibration: at a 95% accuracy threshold, Nimble can auto-handle 43% of traffic versus Qwen's 16%
The takeaway: for production decision-routing, calibration quality — not headline accuracy — determines how much work can be automated, making Nimble 2.7x more effective. Full methodology and setup are linked in the post.
More from Models
- Dev racing to burn $800 of Gemini API credit teases rumored Gemini 4 Argon and Nano Banana Pro 2 — Angaisb_ · 2026-10-09
- 16 VPD weight edits boost model accuracy 10x, revealing the attention head that suppresses introspection — Sauers_ · 2026-10-09
- GPT-6.1 Sol "Ultrafast" clocks in at just ~49 tok/s in user API speed test — RexDouglass · 2026-10-09
- One attention head drives sandbagging-like introspection in Qwen3-1.7B; ablating it helps — Sauers_ · 2026-10-09
- Text-Only Qwen3.5 2B/4B/9B MLX 4-bit Packages Released, 2B Is Just 1GB — sachasayan · 2026-10-09
- Why multilingual LLMs are hard: character decoding and BPE are the hidden bottleneck — ivan_bezdomny · 2026-10-09