Is your LLM eval 71%→74% gain real or noise? Bootstrap CIs explain how to tell

bgoncalves · x · 2026-09-08

A thread from data4sci: your LLM eval says accuracy jumped from 71% to 74% after a prompt change — but is that a real improvement or noise?

Key point: without confidence intervals you can't tell. The thread explains how bootstrap CIs quantify the uncertainty in eval results, helping engineers avoid mistaking random fluctuation for genuine gains when tuning prompts or benchmarking models.

Original post →

More from Research

Research channel →