0.8B model plus 9 LoRA adapters routes agent decisions 38x faster with +8.7 accuracy

Usual_Maximum7673 · reddit · 2026-10-01

A developer released Jeff-Qwen3.5-0.8B, a small "System 1" model that picks between user-defined options with calibrated probabilities in one forward pass, plus 9 task-specific LoRA adapters (40MB each) covering recurring agent decisions: prompt injection detection, tool choice, ticket triage, answer grounding, and more.

The setup: Jeff + adapters answer first and only escalate uncertain queries to Qwen3.8-27B (8-bit, reasoning off). On an M4 Max:

Method rigor: 5 adapters trained mostly on synthetic data with per-row model provenance; every dataset passed a shortcut check (e.g., answers not leaked by text length) and independent review; adapters mixed in 10% of base training data to preserve generality; escalation thresholds tuned on separate calibration rows.

Fully open: weights (Apache 2.0), code (MIT), test and calibration sets published (training data withheld), browser demo available for all nine adapters.

Original post →

More from coding & agent

coding & agent channel →