A deep dive on building frontier-lab evals explains why 100% scores can be a failure
aakashgupta · x · 2026-07-28
A long talk breaks down how to build frontier-lab-quality evals, covering what an eval really is, why old benchmarks stopped working, and how to design better ones.
Key points
- Distinguishes offline evals from production evals.
- Explains why older evals broke down as models saturated Q&A.
- Argues that 100% on a benchmark can mean the benchmark is too easy.
- Walks through how to write an eval, starting from the easiest case and then finding the ceiling.
- Shows a live example of building a real eval and compares models head-to-head.
- Emphasizes that strong evals depend on subject-matter expertise, not just generic prompt craft.
More from Research
- Maker shares first AI robot kit built with Raspberry Pi 5 and Hermes agent — petrusenko_max · 2026-07-28
- Exploring Artificial Life: Wolfram and Others Feature in Lenia Simulation — max_romana · 2026-07-28
- A new artificial-life video asks what’s missing for open-ended evolution — max_romana · 2026-07-28
- Macrocosmos starts a permissionless 16B model training run across three continents — markjeffrey · 2026-07-28
- Free AI curriculum maps a practical path from first principles to LLMs — tetsuoai · 2026-07-28
- Kimi report reveals a wide internal benchmark suite for coding and agent skills — stochasticchasm · 2026-07-28