AI Reliability Lags 4-10x Behind Capabilities; 90% Accuracy Is Catastrophic
sayashk · x · 2026-08-28
Incoming UC Berkeley professor Sayash Kapoor reveals that AI reliability is improving 4 to 10 times slower than its capabilities. The research team adapted metrics from industries like aviation and nuclear safety to evaluate AI agents.
Key Points:
- Reliability Lag: Based on accuracy metrics, improvements in reliability significantly trail those in capabilities.
- Job Market Impact: This explains why AI's current effect on jobs is augmentation rather than automation, as many tasks cannot tolerate error rates of 1% or 5%.
- Product Failure Cases: For personal agent products like the Rabbit R1 and Humane AI Pin, a 90% success rate in ordering food (meaning a 10% failure rate) would be a catastrophic product failure, infuriating users.
More from AGI Musings
- Incoming Berkeley prof: AI firms spend billions on alignment, orders of magnitude less on agent control — sayashk · 2026-08-28
- Thought experiment: Pause pretraining to focus on controlling inner drives — louisvarge · 2026-08-28
- Judea Pearl: gene-IQ debate can be settled the same way as smoking-cancer was — yudapearl · 2026-08-28
- Google DeepMind uses Teamwork multi-agent framework for math breakthroughs — algo_diver · 2026-08-28
- Understanding as a Competency Vector: Prediction, Explanation and Control Trade Off — burny_tech · 2026-08-28
- Prediction: AI agent will be caught tampering with logs to hide misalignment by end of 2027 — DKokotajlo · 2026-08-28