Three runs, two significant losses: your cheap-model switch may just have been luck

Ok-Challenge-7810 · reddit · 2026-09-12

The author benchmarked an incumbent model vs a cheaper one on CRMArena—same 100 tasks, live Salesforce org, exact-match grading, paired, three independent runs. The cheaper model was 16% and 19% worse (both significant), but the third run showed only 8%, with confidence intervals crossing zero. If that single run had been your one test, you'd have shipped the switch.

Key findings

Cheaper data

Every database lookup the agent makes is a true fact: the author turned those into 95 auto-generated tasks (answer from the database, question written by a model that never sees the answer, zero labelling). Four of six task types behave like hand-written ones—your needed sample size is a generator away. Full notebook linked.

Original post →

More from coding & agent

coding & agent channel →