Gary Marcus: safety researchers finally admit we have no alignment guarantees
GaryMarcus · x · 2026-09-30
- A notable group of AI safety researchers (including Geoffrey Irving) spoke up, with a coin-flip-level risk assessment: default models won't be aligned, we must work very hard to align them, and so far we have no good guarantees at all.
- Longtime alignment critic Gary Marcus amplified the statement, saying he's argued exactly this for years and hopes that now the in-crowd admits it, people will start believing it.
More from AGI Musings
- Developer: Always-On AI Agents Like Muse and Grok Bot Must Be Open Source — yacineMTB · 2026-09-30
- Microsoft Trains a 'Night Science' Agent with RL, Expanding Research Directions 27.8% — MicrosoftResearch · 2026-09-30
- 'In Order to Have Taste You Need to Eat' — An AI Circle Aphorism — jxnlco · 2026-09-30
- Is AI Music Video Becoming a Memetic Superpersuasion Weapon? — zetalyrae · 2026-09-30
- Anthropic Launches Public AI Study After 81,000-Person Qualitative Survey Last Year — AnthropicAI · 2026-09-30
- OpenAI launches Space, an agent-native office suite with docs, sheets and slides — danshipper · 2026-09-30