Competence Gate: Using Small Model Confidence for Tool Calls
Synthium- · reddit · 2026-07-05
The author created a 10MB LoRA adapter and orchestration layer for Qwen3.5-4B that decides per query whether to answer directly, search the web, or retrieve local documents, refusing to fabricate facts if unverified. The core idea: small instruction models struggle to verbalize their confidence (tested 7 models from 3–9B, all hitting a confidence ceiling), but their internal activations actually contain this signal. The adapter reads this directly to gate tool calls. Experiments show this gating is better at catching errors than the base model's tool calls (d′ increased by 0.46, 95% CI [0.01,0.89]), and in multi-token cases, 87% were indeed wrong. A dual-signal version routes privacy queries to local retrieval, reducing the proportion of privacy questions sent to public web search from 22% to 10%. It runs locally on Apple Silicon/MLX and offers a GGUF version.
More from coding & agent
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- Steal this idea: prompt-to-hardware where agents assemble custom devices — paraschopra · 2026-09-11
- Model Is the Least Interesting Part: A Guide to Six Core AI Architectures from RAG to Multi-Agent — goyalshaliniuk · 2026-09-11
- Non-coder builds layered memory architecture: 20k tokens tracks a year of agent conversations — matteoianni · 2026-09-11