Taste-Bench: best model flubs 40% of long-horizon agent forks, longer reasoning doesn't help

alex_verem · x · 2026-09-24

A Microsoft/City University of Hong Kong paper, The Tasteful Agent, argues long-run agent outcomes hinge on decision quality at "forks," a capability the authors call taste.

Taste-Bench

Findings

Taste is trainable

A 27B Qwen3.6 trained on fork outcomes gained 17.9 points on unseen forks; advising a coding agent before 41 SWE-bench Pro tasks lifted success from 14.6% to 33.7% — more than double (perfect advice caps at 39%).

The paper directly challenges the "bigger model, longer thinking" pitch: on the decisions that sink long tasks, longer reasoning changes nothing.

Related event: Microsoft's Taste-Bench Shows Top Models Fail 40% of Long-Horizon Decisions(2 posts)→

Original post →

More from coding & agent

coding & agent channel →