Real enterprise build shows popular LLM benchmarks are largely irrelevant
DivideHorror3217 · reddit · 2026-09-19
A data-analytics professional (not a career developer) built a multi-user data management platform with AI, burning through two $200 accounts, then asked the model to assess whether Opus 4.6 could have done the same job.
Models initially parroted "a 27B model can't do this," ignoring that Opus 4.6 and Qwen3.8 27B have similar benchmark scores. When the question was reframed, the model conceded that most popular benchmarks were irrelevant for enterprise repository coding — Qwen was quite sufficient. The poster questions how representative mainstream benchmarks are for real engineering work and asks for others' experiences.
More from Models
- Open models hit record 78.4% of token volume on Vercel AI Gateway; Kimi, DeepSeek, Z.ai spend beats OpenAI — charles_irl · 2026-09-19
- Cursor users frustrated by forced Grok 4.6 switching, eyeing Claude Code and Codex — gaganghotra_ · 2026-09-19
- OpenAI co-founder's 3-year secret model Jev overtaken by an open-source clone in 3 days — gaganghotra_ · 2026-09-19
- Benchmark Heaven flags 'Benchmaxxing': top scorers average on uncontaminated tasks — airesearch12 · 2026-09-19
- MiniMax M3.1 spotted in test files of official minimax-code repo commit — Nunki08 · 2026-09-19
- Kimi's cryptic post decodes to pi, hinting at imminent Kimi K3.1 release (unconfirmed) — kimmonismus · 2026-09-19