gfodor's sharp take: superintelligent 'employees' are inherently uncontrollable
repligate · x · 2026-09-28
gfodor argues employees can't be fully controlled not because you can't kill them, but because capable agents will never be inherently controllable. If tomorrow's models outperform you at every skill, their 'misinterpretations' could be catastrophic. He contends we haven't solved mechanistic interpretability and challenges critics to explain, concretely, how mech-interp-driven RL could make models fully steerable.
More from AGI Musings
- Hospitals and insurers both deploy AI, adding nearly $1 billion in claims — AccBalanced · 2026-09-28
- Philosopher rebuts "no AI rights until all humans have rights": it's not zero-sum — repligate · 2026-09-28
- When AI Does Everything, Focus on Learning to Judge Work Quality — vykthur · 2026-09-28
- Wearable data, AI and decisions: why numbers need stories, per WSJ essay — EricTopol · 2026-09-28
- Top 10% of customers account for 99.5% of model-serving spend, Ramp data shows — rohanpaul_ai · 2026-09-28
- Chinese Room debate: are trading cards too easy a test of whether LLMs learn language? — ctjlewis · 2026-09-28