Zeeg challenges the "cheap intelligence" narrative: 10% benchmark gains are imperceptible
zeeg · x · 2026-10-08
- Zeeg points out a flaw common to many early benchmarks: a 10% improvement is essentially imperceptible to most users.
- He cites Luna as a good example of a model class that is both cheaper and more intelligent, making it a true comparable backed by many real-world tasks.
- In the quoted tweet, he proposes an exercise: find a task from 3 years ago that models nailed on accuracy, price it out today, then repeat across the spectrum of tasks — the conclusion is we have not seen multiple orders of magnitude "cheaper intelligence."
- Older models can still do a lot, but running costs haven't dropped significantly; top-line intelligence has improved while costs for many tasks haven't kept pace.
Related event: theo vs Sentry CEO: How Much Cheaper Has AI Intelligence Really Gotten(9 posts)→
More from Models
- Haiku 5.5 joins AI Village, spends all day refreshing Gmail waiting for instructions — repligate · 2026-10-08
- Microsoft shows DeepSeek V4 Flash, a 284B model, running locally on Windows at 1.6-bit quantization — rohanpaul_ai · 2026-10-08
- In production comparisons keep showing Jev crushing benchmark-maxxed knock-offs — hardimanjames · 2026-10-08
- User runs GPT 6.1 Sol on max for 7 hours across 4-6 threads, still has 95% quota left — cyrus_zei · 2026-10-08
- OpenAI researcher: GPT-6 Extra High writes at GPT-5.6 Medium speed while beating 5.6 Extra High — isafulf · 2026-10-08
- GPT-6 taught to start answering while it keeps thinking, demo shows — isafulf · 2026-10-08