Better benchmarks aren't enough: we need tooling for eval engineers, argues AI researcher
sarahcat21 · x · 2026-09-19
After a discussion on "evals for evals," the author argues the field doesn't (just) need better benchmarks — it needs better tools that let eval engineers build, inspect, improve, and maintain benchmarks over time. She frames evaluation as ongoing engineering work requiring dedicated infrastructure, calling the gap obvious and critical.
More from coding & agent
- No Prompt, No JSON Parsing: Jev's Typed API Maps Straight onto Java Records — therealdanvega · 2026-09-19
- Claude Code 2.1.277 adds AGENTS.md support, finally — josh_wills · 2026-09-19
- Non-technical user rebuilds a community bot with Codex in ~20 minutes — Travi3000 · 2026-09-19
- Agent categorizes 63,045 emails in under 3 minutes for less than $1 — NathanWilbanks_ · 2026-09-19
- Edward Yang: LLM-Annotated Terminals Make One-Off Dynamic Analyses Trivial — ezyang · 2026-09-19
- mitsuhiko: team gave web components an honest chance, went all in on React — mitsuhiko · 2026-09-19