19 fixtures, 6 models, 3 runs: newest AI models didn't beat the old one at finding bugs
ChanceKelch · x · 2026-10-01
Chance Kelch, founder of Flock Synthetics, ran a rigorous eval: 19 curated fixtures, 6 models, 3 runs each, testing how well models interpret recorded evidence to find product bugs. GPT-5.6 Luna scored highest on average, GPT-6 Luna was cheapest, GPT-6 Sol was steadiest — but the oldest model, GPT-4.1, was best on the critical case (an "optional" form field blocking submission). Takeaway: newer isn't better; averages hide tradeoffs and switching models depends on your task and workflow.
More from coding & agent
- New tool makes benchmarking across factory configs trivial, not just models — vikvang1 · 2026-10-01
- Magnitude inference engine hits #1 on HN, claims up to 2x faster local open-model runs than llama.cpp — nickbaumann_ · 2026-10-01
- marimo-lens lets you point at charts and steer your coding agent directly — S_Conradi · 2026-10-01
- Non-coder builds MCP-only medieval trade game played entirely by AI agents — Wainfare · 2026-10-01
- Dev roasts Claude Code's Opus for forgetting UI parameters, makes No Country meme — firasd · 2026-10-01
- Solo fine-tune of Qwen3.8-27B-pi fixes effort ordering, saves 41% tokens — victormustar · 2026-10-01