0.8B model plus 9 LoRA adapters routes agent decisions 38x faster with +8.7 accuracy
Usual_Maximum7673 · reddit · 2026-10-01
A developer released Jeff-Qwen3.5-0.8B, a small "System 1" model that picks between user-defined options with calibrated probabilities in one forward pass, plus 9 task-specific LoRA adapters (40MB each) covering recurring agent decisions: prompt injection detection, tool choice, ticket triage, answer grounding, and more.
The setup: Jeff + adapters answer first and only escalate uncertain queries to Qwen3.8-27B (8-bit, reasoning off). On an M4 Max:
- Accuracy 86.6% → 95.3% (wrong answers 13.4% → 4.7%)
- Time per decision 8.1s → 0.25s (38x faster)
- Memory under 2GB even with all 9 adapters loaded, vs 28.6GB for the 27B
- The five inbox-agent decisions per message: 87.7% → 95.7%, 39x faster; 6 of 9 adapters hit 97-98% on held-out sets
- 27-way emotion classification: 60.6% vs 35.6% at 42x speed; 30ms per decision on GPU
Method rigor: 5 adapters trained mostly on synthetic data with per-row model provenance; every dataset passed a shortcut check (e.g., answers not leaked by text length) and independent review; adapters mixed in 10% of base training data to preserve generality; escalation thresholds tuned on separate calibration rows.
Fully open: weights (Apache 2.0), code (MIT), test and calibration sets published (training data withheld), browser demo available for all nine adapters.
More from coding & agent
- Longtime Dev: "AI Can't Build Maintainable Architectures" Is a 2023 Take — dreamwieber · 2026-10-01
- Economist Paul Novosad: AI-edited code becomes unmanageable, so I wall it off — paulnovosad · 2026-10-01
- Local models can now power computer use agents, but regulated industries still lack a playbook — Ambitious_Fold_2874 · 2026-10-01
- Xiaomi MiMo-V2.6 report: one mixed GRPO run across all domains, $2.6M for Pro — SergioPaniego · 2026-10-01
- Synentra: an open-source gateway that lets AI agents' actions be judged by intent and risk before execution — ziagham · 2026-10-01
- Let agents write code to call tools instead of exposing every tool directly — Future_AGI · 2026-10-01