Benchmarks are meaningless, the differences between models are just vibes, dev argues
gnukeith · x · 2026-09-22
Developer gnukeith rants that benchmarks are "absolutely meaningless," that the scale-of-intelligence narrative is pure hype and marketing, and that the only real difference between models is "vibe" — while joking that nobody, except perhaps Google itself, knows what Google is doing. The take captures growing fatigue with leaderboard-driven model evaluation.
Related event: Dev claims benchmarks are meaningless, only 'vibes' separate models(2 posts)→
More from Models
- Anthropic investigates elevated errors across Claude Mythos 5.1, Fable 5.1 and Opus 5 — ClaudeAI-mod-bot · 2026-09-22
- Dev Swaps Opus for Mimo-v2.6 in Cline on Client Projects: 'It's a Beast' — MicahBerkley · 2026-09-22
- Kev refactored onto Qwen3.5: open-source decision models now at 0.8B, 4B and 9B — alexcovo_eth · 2026-09-22
- Why OpenAI bets on math: it's the most verifiable domain for reinforcement learning — burny_tech · 2026-09-22
- Open-source decision model Laya ported to Core ML: 99.5% ops on ANE, 3.7ms per decision — alexcovo_eth · 2026-09-22
- Fireworks: routing 18 models per task hits 97.6% solve rate at $1.88 vs best single model's 74.1% at $6.52 — sophiamyang · 2026-09-22