When should a fast agent defer? Deferring least-confident 30% to a reasoner gains up to +0.13
Gian Luca Bailo · hf · 2026-10-09
System Switch studies when a fast decision model in a dual-process agent should hand control to a slow reasoning VLM, while the game keeps running. Built on closed-loop Doom and the new open "System One" typed-decision models (0.15B-9B), served via llama.cpp.
Findings on 900 held-out questions:
- Zero-shot decision models pick "collect items" 1.6-1.8x more often than chance among their errors, regardless of option order
- Accuracy, calibration, and sensitivity are distinct: models with similar accuracy differ widely in AUROC
- Offline: deferring the least-confident 30% of decisions gains in proportion to the actor's AUROC (rank correlation 0.87); actual gains +0.13 (original option order) / +0.08 (shuffled), reasoning carrying about half
- In closed loop (33 games, 3 seeds), no variant reaches the exit; committing to plans (the reasoner's or a fixed explore rule) opens more doors, but the reasoner mistakes ordinary doors for locked ones when told some doors need keys
Code, prompts, data, and logs are released.
More from Research
- LLMs reportedly advance on 4 of 7 Millennium Prize Problems, 3hrs compute each — ycombinator · 2026-10-09
- New research: conflicting training values can make models' CoT contradict their answers — OwainEvans_UK · 2026-10-09
- Arena raises $200M Series B at $3.1B valuation, launches agent Alignment Index — a16z · 2026-10-09
- PAMI anchors object motion to body parts for text-to-HOI, +14.5% contact recall — Chuqiao Li · 2026-10-09
- Δ-MOPD transfers teacher shifts instead of endpoints, +4.11 Math in multi-teacher distillation — LinkedIn · 2026-10-09
- Lineage-gated agent memory wipes 18.8-25.5% cross-department leakage at 13.8μs overhead — Venkata M Sangaraju · 2026-10-09