Dev says model benchmarks are broken: real prompts run an hour+, not $1.50 tasks

pvncher · x · 2026-09-27

pvncher argues model benchmarks are 'completely broken' and don't reflect real-world use. His point: benchmarks like deepswe test tiny isolated tasks that cost about $1.50 of inference, with clear upfront requirements. Real users run prompts that take an hour+ — the model must decipher ambiguous instructions, navigate a messy repo, and iterate with compactions and human steering along the way. He says stop putting so much trust in these scores.

Original post →

More from Models

Models channel →