Box^2-Bench Shows Frontier Models Struggle to Reject Unreliable Workflow Guidance
Minghan Wang · hf · 2026-10-01
Box^2-Bench varies workflow reliability while fixing model and task, measuring whether models can selectively rely on guidance. Frontier models benefit from reliable workflows but stay vulnerable to misleading ones. Counterfactual SFT improves robustness while outcome-based RL shifts the balance toward helpful workflows, and the behavior extends to peer correction and corrupted-memory robustness.
More from coding & agent
- Hybris MCP Server lets AI assistants manage SAP Commerce Cloud instances — modelcontextprotocol · 2026-10-01
- cua-speedrun: CMU benchmark shows 4.4x speed gap between equal-scoring computer-use agents — arankomatsuzaki · 2026-10-01
- M-Anchor: a deterministic, zero-LLM Python gate that blocks unsupported LLM record updates — Informal-Winter-3190 · 2026-10-01
- Dev Burns Through Claude Weekly Quota in Under Two Days on Opus 4.5 Alone — yihui_indie · 2026-10-01
- Split coding workflow: Opus 5.5 xhigh for planning, Sol 6.1 high for the bulk of coding — haider1 · 2026-10-01
- FV expert: AI folks' view of formal verification is 15 years out of date — tianyin_xu · 2026-10-01