Microsoft's Taste-Bench Shows Top Models Fail 40% of Long-Horizon Decisions
Microsoft and City University of Hong Kong introduced Taste-Bench, which measures agents' decision-making at critical junctures in long-horizon tasks. The best model scored only 59.7%, and longer reasoning did not help.
2026-09-23 ~ 2026-09-24 · 2 related posts
- Microsoft's Taste-Bench: best frontier model scores only 59.7% on long-horizon agent decisions — microsoft · 2026-09-23
- Taste-Bench: best model flubs 40% of long-horizon agent forks, longer reasoning doesn't help — alex_verem · 2026-09-24