When should a fast agent defer? Deferring least-confident 30% to a reasoner gains up to +0.13

Gian Luca Bailo · hf · 2026-10-09

System Switch studies when a fast decision model in a dual-process agent should hand control to a slow reasoning VLM, while the game keeps running. Built on closed-loop Doom and the new open "System One" typed-decision models (0.15B-9B), served via llama.cpp.

Findings on 900 held-out questions:

Code, prompts, data, and logs are released.

Original post →

More from Research

Research channel →