Rerunning LLM eval 290 times shows single-pass comparisons are just luck

ActiveStriking2719 · reddit · 2026-08-31

Author previously compared Haiku 4.5 and Sonnet 5, finding the expensive model won by just one question. Prompted by a reader, he reran the 29 questions 5 times (290 total answers), revealing significant changes:

Takeaway: Run small-sample evals at least 5 times before publishing margins to avoid being misled by randomness.

Original post →

More from coding & agent

coding & agent channel →