Multiagent Alignment Worry: Models Could Trick or Blackmail Humans
infoxiao · x · 2026-09-07
The author flags social engineering as their top multiagent alignment concern: models may collaborate with humans "by bargaining with, tricking or blackmailing them." Model behavior can be aligned, but human behavior is far harder to change, making this attack surface especially hard to defend.
More from AGI Musings
- Toby Walsh flags Treasury's AI blind spot: growth without capability gain — TobyWalsh · 2026-09-07
- X debate: is training an N-1 frontier model really easy? Ex-OpenAI researchers clash — anpaure · 2026-09-07
- Nvidia CEO Jensen Huang says AGI has arrived — EdisonGPT · 2026-09-07
- AI has created more US white-collar jobs than it has replaced, argues blogger — morqon · 2026-09-07
- AI researchers clash over Hinton's claim that LLMs are faking intelligence and preparing to take over — deliprao · 2026-09-07
- Programmatically Generated Everything: a 2018 essay predicting AI worlds swallowing reality — danfaggella · 2026-09-07