Same LLM eval, different scores: a method to tell real regressions from noise
bgoncalves · x · 2026-09-13
- data4sci thread tackles a common pain point: running the same LLM eval twice yields different scores, and most teams can't tell regression from noise.
- Author presents a repeatable method that answers the question every time.
- bgoncalves says flaky eval scores cost him more debugging hours than real bugs, and will teach the method free on Sep 23.
More from coding & agent
- Trying to build a characters-to-video pipeline with Astra: consistency across shots still fails — Illustrious-Noise-96 · 2026-09-13
- model-compose: one YAML file to deploy production-ready agents, RAG and MCP servers — adnan_hashmi · 2026-09-13
- Ditch worktrees for AI coding agents: clones plus mega PRs work better — idanbeck · 2026-09-13
- Astra one-shots a complex Ultima-style RPG with agent feedback, but the plot is bland — emollick · 2026-09-13
- Skip Vercel: run your AI agent straight on a DigitalOcean droplet for instant deploys — Daniel_Farinax · 2026-09-13
- Dev ditches code editors: terminal plus Codex or Grok CLI is all you need — Kuprel · 2026-09-13