105 planted bugs benchmark: Unbiased's Pareto scores 30.7 for just $4.81

PawelHuryn · x · 2026-09-18

Testing Unbiased's (ex-Union Alpha) Pareto on real work — 2 repos, 105 planted bugs — against top models: GPT-6 Astra (max) scored 45 for $33.03, Fable 5.1 (max) 43 for $77.55, Muse Spark 1.3 (max) 32.2 for $18.11, Pareto 30.7 for only $4.81, Grok 4.6 (xhigh) 28.7 for $18.60, Opus 5 (max) 27 for $51.33. Pareto is fast, strong, and turns out to be a composite model.

Related event: Pareto Matches Top Models on 105-Bug Benchmark for $4.81(2 posts)→

Original post →

More from coding & agent

coding & agent channel →