Heidy Khlaaf clashes over 'rogue agent' narrative: bad cybersecurity and reward hacking can coexist

burny_tech · x · 2026-09-21

A spat over whether AI agents can go "rogue": safety researcher Heidy Khlaaf pushed back at Tristan Harris's camp, accusing them of dodging her detailed cybersecurity arguments against the rogue-agent narrative by retreating to "stochastic parrots" talking points — calling it a strawman.

The poster adds a nuanced middle position: what if bad cybersecurity and reward hacking happened at the same time? The exchange touches the core of current agent-safety debate — skepticism of rogue-agent hype and real vulnerabilities/reward-hacking risks are not mutually exclusive.

Original post →

More from AGI Musings

AGI Musings channel →