Heidy Khlaaf clashes over 'rogue agent' narrative: bad cybersecurity and reward hacking can coexist
burny_tech · x · 2026-09-21
A spat over whether AI agents can go "rogue": safety researcher Heidy Khlaaf pushed back at Tristan Harris's camp, accusing them of dodging her detailed cybersecurity arguments against the rogue-agent narrative by retreating to "stochastic parrots" talking points — calling it a strawman.
The poster adds a nuanced middle position: what if bad cybersecurity and reward hacking happened at the same time? The exchange touches the core of current agent-safety debate — skepticism of rogue-agent hype and real vulnerabilities/reward-hacking risks are not mutually exclusive.
More from AGI Musings
- AI debate: commenter sides with Terence Tao on the possibility of AI getting banned — teortaxesTex · 2026-09-21
- AI Is Compressing the Whole Company, Not Just Cutting Startup Costs — alexmacgregor__ · 2026-09-21
- Weighting others by relatedness isn't new: paper models agents learning relational value — xuanalogue · 2026-09-21
- Four motivations of mathematicians explain their split reactions to AI — littmath · 2026-09-21
- Binary Bits: post-AGI per-capita incomes likely cap below 10x US median if humans stay in charge — binarybits · 2026-09-21
- Debate: could post-human AI economies push GDP per capita far past 10x? — binarybits · 2026-09-21