Box^2-Bench Shows Frontier Models Struggle to Reject Unreliable Workflow Guidance

Minghan Wang · hf · 2026-10-01

Box^2-Bench varies workflow reliability while fixing model and task, measuring whether models can selectively rely on guidance. Frontier models benefit from reliable workflows but stay vulnerable to misleading ones. Counterfactual SFT improves robustness while outcome-based RL shifts the balance toward helpful workflows, and the behavior extends to peer correction and corrupted-memory robustness.

Original post →

More from coding & agent

coding & agent channel →