"Trust the price, not the benchmarks": signull argues model pricing reveals real capability
signulll · x · 2026-10-01
- Deedy laments that frontier models must win a pile of benchmarks before launch, calling the numbers meaningless.
- signull counters that price is the only trustworthy signal: high pricing signals genuine confidence, while cheap high scores suggest "benchmaxxing."
- The thread devolves into jokes about a "felonybench" measuring how many sites you can hack into.
Takeaway: vendor pricing strategy may be a more honest capability signal than leaderboard scores prone to overfitting.
More from Models
- Can local LLMs handle Blender and game dev? Reddit says they still fall short — Any-Lingonberry7411 · 2026-10-01
- Rox benchmarks: Jev reranking beats GPT-5 Mini — 20x faster, 10x cheaper, 12% more accurate — hardimanjames · 2026-10-01
- Nat Lambert: More Frontier Labs Like Google's Gemini 4 Benefit Consumers — natolambert · 2026-10-01
- Qwen3.8-Flash-Next Cut 44% via REAP Hits 70% on Terminal-Bench 2.1 — rmonsurate · 2026-10-01
- Google launches Gemini 4 Argon, a cybersecurity model that tops prompt injection benchmarks — ralucaadapopa · 2026-10-01
- Model wars: OpenAI went from best model in the world to arguably third place in a week — signulll · 2026-10-01