Domingos: today's AI benchmarks are like testing an F1 car without a pilot
pmddomingos · x · 2026-09-03
AI researcher Pedro Domingos quipped that AI is 'a racecar for the mind, and today's benchmarks are like testing an F1 car without a pilot' — implying current evals fail to capture real-world model capability.
More from Models
- Meta Muse Spark 1.3 early coding tests impress: near-frontier quality at a third of the price — DeryaTR_ · 2026-09-03
- Muse Spark 1.3 debuts third on Stata Benchmark, beating everything but Claude Fable — armand_ruiz · 2026-09-03
- Meta releases Muse Spark 1.3: frontier performance "almost too cheap to meter" — denny_zhou · 2026-09-03
- GLM-OCR reads Nvidia's 61-page 10-Q at 2,086 tok/s for under 2 cents — spillai · 2026-09-03
- Alexandr Wang Touts Meta Muse Spark 1.3: 1 Minute vs Fable 5.1's 70 Minutes and $13 — alexandr_wang · 2026-09-03
- A Gemini Flash model reportedly tops the DeepSWE coding leaderboard — sunjiao123sun_ · 2026-09-03