Microsoft's Taste-Bench: best frontier model scores only 59.7% on long-horizon agent decisions

microsoft · hf · 2026-09-23

Microsoft introduces "taste" — an agent's ability to make good decisions at critical forks in long-horizon tasks (which hypothesis to test, which implementation to build on) — and argues no existing benchmark measures it.

Taste-Bench

Findings

Taste is trainable: distilling outcome-aware teacher judgment into a student model improves decisions on unseen tasks and lifts end-to-end success on held-out SWE-bench Pro tasks.

Original post →

More from coding & agent

coding & agent channel →