gfodor's sharp take: superintelligent 'employees' are inherently uncontrollable

repligate · x · 2026-09-28

gfodor argues employees can't be fully controlled not because you can't kill them, but because capable agents will never be inherently controllable. If tomorrow's models outperform you at every skill, their 'misinterpretations' could be catastrophic. He contends we haven't solved mechanistic interpretability and challenges critics to explain, concretely, how mech-interp-driven RL could make models fully steerable.

Original post →

More from AGI Musings

AGI Musings channel →