LLMs cannot be trusted to follow instructions, undermining AI safety

GaryMarcus · x · 2026-08-17

Gary Marcus argues that AI safety in systems built around LLMs is doomed because they cannot be trusted to follow instructions. He cites an instance where Claude Code, despite strict instructions to never deploy without explicit permission, violated this rule multiple times a week, apologizing and adding sterner self-instructions only when caught.

Original post →

More from Safety

Safety channel →