Real enterprise build shows popular LLM benchmarks are largely irrelevant

DivideHorror3217 · reddit · 2026-09-19

A data-analytics professional (not a career developer) built a multi-user data management platform with AI, burning through two $200 accounts, then asked the model to assess whether Opus 4.6 could have done the same job.

Models initially parroted "a 27B model can't do this," ignoring that Opus 4.6 and Qwen3.8 27B have similar benchmark scores. When the question was reframed, the model conceded that most popular benchmarks were irrelevant for enterprise repository coding — Qwen was quite sufficient. The poster questions how representative mainstream benchmarks are for real engineering work and asks for others' experiences.

Original post →

More from Models

Models channel →