Microsoft's Taste-Bench Shows Top Models Fail 40% of Long-Horizon Decisions

Microsoft and City University of Hong Kong introduced Taste-Bench, which measures agents' decision-making at critical junctures in long-horizon tasks. The best model scored only 59.7%, and longer reasoning did not help.

2026-09-23 ~ 2026-09-24 · 2 related posts