Microsoft dev blog: building AX evals that actually work, when scalar scores mislead
lee_stott · x · 2026-09-07
Microsoft Principal Developer Advocate Waldek Mastykarz wraps up his Agent Experience (AX) series with a guide to building agent evals that actually work. Most evals, he argues, produce confident, consistent, and meaningless results: contaminated data, scenarios that don't reflect real usage, criteria checking the wrong thing, and scores that rise while developer experience stays flat.
Key points:
- Evals are a means to an end: find where models have gaps using your tech, then fill those gaps.
- Generated instructions (have a strong model write instructions for a cheaper one) work surprisingly often but small models have limits — the only way to know is evals.
- Prior articles covered what to measure, why benchmarks don't transfer, and hidden variables.
More from coding & agent
- AI coding instruction files grow 226% on average; study proposes 'catastrophic remembering' and comments as fix — rohanpaul_ai · 2026-09-07
- Paper: AI coding prompts grow 226% due to 'catastrophic remembering'; comments can cut excess by 99.3% — rohanpaul_ai · 2026-09-07
- Claude Code Users Press Anthropic on Whether Usage Resets Are Still Manual or Automatic — burhop · 2026-09-07
- Devtoolsniff bundles cursor rules generators, MCP configs, and an AI IDE cost calculator — Ice-Medium · 2026-09-07
- Blogger says Google Astra's natural tone broke his last dependency on Claude — StewartalsopIII · 2026-09-07
- AI agent sorts 15,000 files, 500GB Google Drive in 15 minutes with zero deletes — thisiskp_ · 2026-09-07