What do PR maxxers do when your test suite costs $10 to run?
zeeg · x · 2026-10-07
David Cramer poses the practical question behind the debate: what do you do when your agent test suite costs $10 per run due to real LLM inference? It kicks off the discussion on mocks vs real-model testing.
More from coding & agent
- GitHub's busiest repo saw ~1B requests a month; storage layer rebuilt for agent-scale dev — marlene_zw · 2026-10-07
- Dev slams Linear for shipping "just a chat box" instead of rethinking work for AI agents — devenbhooshan · 2026-10-07
- Dev claims 700,000 lines of code in 4 days — "just me and Opus" — haydendevs · 2026-10-07
- Opus 5.5 nearly matches Fable 5.1 at two-thirds the cost in real-repo bug benchmark — PawelHuryn · 2026-10-07
- Bug Hunt Benchmark yields different rankings for frontier models — PawelHuryn · 2026-10-07
- DAEDALUS bootstraps agent memory from self-generated tasks, +15.9 points success rate — illuin · 2026-10-07