Stop chasing phantom LLM eval swings: bootstrap confidence intervals in 20 lines of Python

bgoncalves · x · 2026-09-23

Core idea

You ran the same eval twice and got 84.2, then 81.9. Without error bars, every score is a coin flip dressed as a fact — teams chase phantom regressions and burn weeks on noise.

The method

Bruno Gonçalves (author of the Data For Science newsletter, former NYU Data Science fellow) will teach in a free 30-minute live session (recorded):

Original post →

More from coding & agent

coding & agent channel →