Microsoft dev blog: building AX evals that actually work, when scalar scores mislead

lee_stott · x · 2026-09-07

Microsoft Principal Developer Advocate Waldek Mastykarz wraps up his Agent Experience (AX) series with a guide to building agent evals that actually work. Most evals, he argues, produce confident, consistent, and meaningless results: contaminated data, scenarios that don't reflect real usage, criteria checking the wrong thing, and scores that rise while developer experience stays flat.

Key points:

Original post →

More from coding & agent

coding & agent channel →