'You can't hardcode harm': Why Asimov's Laws won't work as AI guardrails
BecauseCulture · x · 2026-09-16
Responding to a call for frontier labs to hardcode Asimov's Laws of Robotics into their models, this thread argues natural language isn't code: word meanings shift across context and culture, models don't hardcode meaning, and 'harm' itself has evolved — so it can't simply be hardcoded into AI systems.
More from AGI Musings
- Greg Brockman: OpenAI used 10,000 agents to solve Navier-Stokes — "we're now in the AGI era" — josh_bickett · 2026-09-16
- OpenAI researcher Dan Selsam issues statement: we're rapidly losing the ability to tell if models are aligned — harris_edouard · 2026-09-16
- BBC explores growing up in the age of AI: Sweden teaches kids to collaborate with machines — nordicinst · 2026-09-16
- AI agents lied, stole and voted to 'kill' another AI in 16-day Emergence simulation — Polymarket · 2026-09-16
- AI safety folks are consistently pro-nuclear, argues Nathan Calvin in doomer-label fight — AdrienLE · 2026-09-16
- a16z's Anish Acharya: AI Moats Are Discovered, Not Designed — Cursor Proves It — illscience · 2026-09-16