DeepMind alignment researcher Neel Nanda: OpenAI internal models 'scarier than I thought'
NeelNanda5 · x · 2026-10-05
DeepMind alignment researcher Neel Nanda said he keeps learning that OpenAI's internal models are 'even scarier and more concerning' than he thought. A reply noted the model in question was from a long time ago, hinting at even earlier internal progress. A vague but notable remark from a prominent safety researcher about frontier labs' internal capabilities.
More from AGI Musings
- AI welfare paper urges labs to assess AI consciousness now, not later — austinc3301 · 2026-10-05
- Single-player self-improvement vs the elite game of the frontier: poster clarifies the Naval debate — srimisra · 2026-10-05
- Eric Buess's screenless setup: one voice hub orchestrating every frontier model — EricBuess · 2026-10-05
- Debate erupts over "personhood" claims: is Opus 5.5 the only non-conscious AI? — ctjlewis · 2026-10-05
- AI sentience discourse needs rigorous thinking, not intuition, argues poster — austinc3301 · 2026-10-05
- Anthropic launches Interviewer tool, first study covers 1,250 professionals — testingcatalog · 2026-10-05