Investor Deedy Says AI Benchmarks Are Gamed—Trust Pricing Instead
Silicon Valley investor Deedy Das (@deedydas) posted a thread on October 1 making a striking claim: frontier AI models have entered a stage of "late model capitalism"—you have to "win" on a pile of benchmarks before release just to qualify, and these scores have largely lost their reference value. He also revealed that almost every model company pays benchmark evaluation agencies to pre-test their models (and multiple versions of them) ahead of release, which is precisely how these evaluation firms make money.
Confirmed
- Deedy explicitly stated that benchmark scores are not trustworthy; the truly credible signal is pricing: a vendor daring to charge a high price suggests the model is genuinely strong, while low prices paired with high scores usually mean "benchmaxx" (optimized just for benchmarks).
- He added that the specific price matters less than the signal of being "within ±25% of the frontier range."
- signull (@signulll) echoed the view, arguing that high pricing reflects the vendor's own confidence.
Why it matters
- The claim points directly at a conflict of interest between evaluation agencies and model vendors: if evaluation firms mainly profit from paid pre-release testing by vendors, the independence of their public leaderboards is questionable.
- The discussion (including Deedy and signull riffing on the "felonybench is the only real leaderboard" meme, with signull joking that it measures "how many websites you can hack into") reflects the AI community's collective distrust of the current evaluation system, and offers observers an alternative signal—inferring model strength from pricing.
2026-10-01 ~ 2026-10-01 · 5 related posts
Primary sources
- Deedy: model companies pay benchmark firms to pre-test models — price is the only honest signal — deedydas ·
- Deedy: Trust Pricing, Not Benchmarks — High Price Means a Genuinely Strong Frontier Model — deedydas ·
- "Trust the price, not the benchmarks": signull argues model pricing reveals real capability — signulll ·
- [source] Deedy: Trust Pricing, Not Benchmarks — High Price Means a Genuinely Strong Frontier Model — deedydas · 2026-10-01
- [source] Deedy: model companies pay benchmark firms to pre-test models — price is the only honest signal — deedydas · 2026-10-01
- [source] "Trust the price, not the benchmarks": signull argues model pricing reveals real capability — signulll · 2026-10-01
- "felonybench is the one true bench": AI circle jokes about benchmark skepticism — deedydas · 2026-10-01
- Investor: Model companies pay benchmark firms to test multiple versions pre-launch — deedydas · 2026-10-01