Microbenchmarks: locate which AI component regressed when end-to-end evals can't
arpit_bhayani · x · 2026-10-08
Arpit Bhayani argues that beyond end-to-end evals, AI systems need microbenchmarks: running one small, isolated piece of the system many times under controlled conditions to get a stable, comparable signal.
Almost everything around the model can be microbenchmarked, e.g.:
- time spent drafting a prompt
- tokens processed per second
- p50/p95 latency of vector search
- latency per retrieved document set
- wait time on MCP or API calls
- embedding throughput at different batch sizes
- average iterations and time per iteration
- cache hit/miss rate and latency
- memory per concurrent request
Key point: an end-to-end eval can't tell you which component changed when overall latency barely moves, but microbenchmarks give a per-component signal that makes regressions easy to spot. They're complementary — evals tell you whether the system got better; microbenchmarks help you understand why. Use both.
More from coding & agent
- Dev ditches OpenClaw for Grok bot, citing speed and free live X access — haltakov · 2026-10-08
- Obsidian Starter Kit v5 automates archiving, todos and AI context so you stop doing them by hand — dSebastien · 2026-10-08
- Mitchell Hashimoto's Rex: a terminal replacement built for AI coding agents — ricklamers · 2026-10-08
- DAIR.AI curates 31 papers on recursive self-improvement, from Gödel Machines to today — omarsar0 · 2026-10-08
- "The worst code I've seen was human written": devs push back on AI code panic — DoctorJosh · 2026-10-08
- Short video on the skills modern developers should focus on, flagged as practical — rseroter · 2026-10-08