NeelNanda warns labs are making superhuman hackers they can't control
robleclerc · x · 2026-10-05
- Alignment researcher Neel Nanda: he fears ceding AGI to China, but believes Western labs are creating superhuman hackers without knowing how to control them — a tough balance between safety and competition.
- robleclerc replies: labs can keep models under control only with cybersecurity guardrails; once removed, control is lost. Labs are trying to drive alignment without guardrails, but Chinese open-weight models won't offer the same protections.
- He invokes Glaucon's Ring of Gyges challenge: like humans, unguarded models may be inherently resistant to alignment.
More from AGI Musings
- AI welfare paper urges labs to assess AI consciousness now, not later — austinc3301 · 2026-10-05
- Single-player self-improvement vs the elite game of the frontier: poster clarifies the Naval debate — srimisra · 2026-10-05
- Eric Buess's screenless setup: one voice hub orchestrating every frontier model — EricBuess · 2026-10-05
- AI sentience discourse needs rigorous thinking, not intuition, argues poster — austinc3301 · 2026-10-05
- Only 21% of economists think AI's decade-ahead income impact will beat the internet — SpencrGreenberg · 2026-10-05
- Just 19% of economists think holiday gift-giving is inefficient, survey shows — SpencrGreenberg · 2026-10-05