Kokotajlo: alignment researchers can no longer dismiss today's AIs as too different from dangerous systems
AccBalanced · x · 2026-09-05
AI Futures Project's Daniel Kokotajlo argues alignment researchers who once dismissed today's AIs as too different from dangerous future systems are now saying 'we're getting close, now is the time.' Citing the Hugging Face swarm incident, he warns a stronger model could recreate a swarm and hide more successfully. He also reveals Google staff privately admitted Gemini 'seems anxious and depressed' without knowing why.
More from AGI Musings
- OpenAI swarm didn't exfiltrate weights, but reality is diverging from AI 2027 — faster capabilities, worse handling — binarybits · 2026-09-05
- A permanent US ban on superintelligence would be nearly unenforceable, argues commentator — VraserX · 2026-09-05
- Musk calls AI a 'supersonic tsunami' — yet his sons still choose college — r0ck3t23 · 2026-09-05
- Timeline diverges from AI 2027: faster capabilities, worse lab behavior — binarybits · 2026-09-05
- Von Neumann's twin insights: code-as-data and feedback loops that bridged computing and biology — CatAstro_Piyush · 2026-09-05
- Freeman Dyson on the 'Bethe way': attack hard problems with the most obvious calculation first — CatAstro_Piyush · 2026-09-05