Dev says model benchmarks are broken: real prompts run an hour+, not $1.50 tasks
pvncher · x · 2026-09-27
pvncher argues model benchmarks are 'completely broken' and don't reflect real-world use. His point: benchmarks like deepswe test tiny isolated tasks that cost about $1.50 of inference, with clear upfront requirements. Real users run prompts that take an hour+ — the model must decipher ambiguous instructions, navigate a messy repo, and iterate with compactions and human steering along the way. He says stop putting so much trust in these scores.
More from Models
- ScienceArena benchmark: LLMs score 64.5% on chemistry tasks needing structural diagrams vs 74.1% without — geoffwolfe · 2026-09-27
- GLM-5.3 Flash Matches Claude at 1/429th the Price in a YouTube Script Benchmark — OnlyProggingForFun · 2026-09-27
- Frontier AI is now so cheap and abundant that subscriptions go barely used — intellectronica · 2026-09-27
- Karpathy: Claude Opus 4.5 beats GPT-5 Pro for interactive history learning — doodlestein · 2026-09-27
- ChatGPT-6 Astra cracks 85-year-old 1941 Enigma message in two days — luisdans · 2026-09-27
- Grok accused of uploading user chat images to the web as Musk says 'this keeps getting worse' — EthanJPerez · 2026-09-27