Microsoft's Taste-Bench: frontier agents pick the better direction only ~60% of the time
omarsar0 · x · 2026-09-25
A Microsoft-led paper introduces Taste-Bench, which tests whether agents have good "taste" at decision forks in long tasks. Forks are auto-mined from parallel attempts and detours in engineering/research runs. The best model answers just 59.7% correctly; forks with late-appearing evidence are much harder, and a larger reasoning budget doesn't help. Distilling a teacher's outcome judgment into a student model improves end-to-end success on held-out tasks.
More from coding & agent
- pnpm urges devs to spend spare tokens fixing its 629 open issues, ships a ready-made agent prompt — itsOmSarraf_ · 2026-09-25
- Geoffrey Huntley Demos a Software Factory Where the Product Is Its Own IDE — teropa · 2026-09-25
- Open-Source iCloud MCP Runs Without a Mac: 31 Tools, Headless Chromium — sjdonado · 2026-09-25
- Microsoft ships enterprise product on OpenClaw after months of joint hardening work — steipete · 2026-09-25
- Open-sourced skill turns your codebase into a polished product promo video via Claude Code — op7418 · 2026-09-25
- BlackRock paper sparks debate: agent payments settle fine, but revoked permissions can't catch up — tallmetommy · 2026-09-25