Microsoft's Taste-Bench: frontier agents pick the better direction only ~60% of the time

omarsar0 · x · 2026-09-25

A Microsoft-led paper introduces Taste-Bench, which tests whether agents have good "taste" at decision forks in long tasks. Forks are auto-mined from parallel attempts and detours in engineering/research runs. The best model answers just 59.7% correctly; forks with late-appearing evidence are much harder, and a larger reasoning budget doesn't help. Distilling a teacher's outcome judgment into a student model improves end-to-end success on held-out tasks.

Original post →

More from coding & agent

coding & agent channel →