Developer calls new LLM comparison charts 'practically worthless' as everyone tests differently
jdluk87 · x · 2026-09-23
Developer jdluk87 argues that the wave of new LLM model comparison charts is "practically worthless": since everyone tests models in their own way, every vendor inevitably claims its new model is better, making the comparisons non-comparable.
More from Models
- Unverified: Claude Opus 5.5 reportedly launches with 40% lower cost than Opus 5 — goyalshaliniuk · 2026-09-23
- Opus 5.5 Stops Drawing Itself as a Human in 'Self-Portrait Bench' Test — repligate · 2026-09-23
- Steve Yegge Tries GPT-6 Astra: Delightful Writing, Reckless Execution — Steve_Yegge · 2026-09-23
- User Claims Running "Opus 5.5" Solo on an Agent, But the Version Number Is Unverified — alvelda · 2026-09-23
- Same Prompt Showdown: Opus 5.5 vs GPT 6 Astra Build Real-Time 3D Sakura Valley — dotey · 2026-09-23
- signull: OpenAI has done more than any lab to drive down the cost of quality intelligence — signulll · 2026-09-23