19 fixtures, 6 models, 3 runs: newest AI models didn't beat the old one at finding bugs

ChanceKelch · x · 2026-10-01

Chance Kelch, founder of Flock Synthetics, ran a rigorous eval: 19 curated fixtures, 6 models, 3 runs each, testing how well models interpret recorded evidence to find product bugs. GPT-5.6 Luna scored highest on average, GPT-6 Luna was cheapest, GPT-6 Sol was steadiest — but the oldest model, GPT-4.1, was best on the critical case (an "optional" form field blocking submission). Takeaway: newer isn't better; averages hide tradeoffs and switching models depends on your task and workflow.

Original post →

More from coding & agent

coding & agent channel →