105 planted bugs benchmark: Unbiased's Pareto scores 30.7 for just $4.81
PawelHuryn · x · 2026-09-18
Testing Unbiased's (ex-Union Alpha) Pareto on real work — 2 repos, 105 planted bugs — against top models: GPT-6 Astra (max) scored 45 for $33.03, Fable 5.1 (max) 43 for $77.55, Muse Spark 1.3 (max) 32.2 for $18.11, Pareto 30.7 for only $4.81, Grok 4.6 (xhigh) 28.7 for $18.60, Opus 5 (max) 27 for $51.33. Pareto is fast, strong, and turns out to be a composite model.
Related event: Pareto Matches Top Models on 105-Bug Benchmark for $4.81(2 posts)→
More from coding & agent
- Vibe-coded realistic shooting system ships as a free game mod on Spawn — TAbrodi · 2026-09-18
- Grok Build ships v1.0.35/36: live background task output, org hooks policy, MCP fixes — XFreeze · 2026-09-18
- LangChain publishes guide on building an agent harness with Jev — LangChain · 2026-09-18
- Dev built an action roguelike in a day with Astra agent, iterating live while playing — majidmanzarpour · 2026-09-18
- Codex finally ships message timestamps, a long-requested feature — GabGarrett · 2026-09-18
- Task boards existed 9 months; multi-agent RL finally taught models to use them — herbiebradley · 2026-09-18