AI safety's real bottleneck: fixing incentive and responsibility chains across the stack
joshua_saxe · x · 2026-10-10
Joshua Saxe argues correct credit assignment across the full chain of responsibility is the highest-order problem in AI safety: bad contractor data → AI lab data team → model errors → legal-AI vendor → unsupervised agent use at a law firm → clients losing cases. Each node needs matching penalties and feedback to update. Today AI companies market weekend-long autonomy without bearing corresponding responsibility, share too little incident information, and spend little on safety — fixing these feedback loops, he says, must happen before AI automation spreads through the economy.
More from AGI Musings
- Dogs Hit Their Own Singularity 20,000 Years Ago — and a Billion Dogs Now Never Need to Work — inductionheads · 2026-10-10
- Pattern matching or inductive bias? Fleuret and syhw spar over what deep learning really is — syhw · 2026-10-10
- Debate revisits Yudkowsky's That Alien Message: physics' low Kolmogorov complexity means AI could locate dangerous tech fast — jd_pressman · 2026-10-10
- If Claude were truly conscious, it wouldn't give itself a 15% chance of being so — inductionheads · 2026-10-10
- Model welfare is a ridiculous hill to die on while human suffering persists, says wolfie_ — emax · 2026-10-10
- Does heavy RL training break the 'LLMs are a blurry upload of humanity' intuition? — jd_pressman · 2026-10-10