LLMs cannot be trusted to follow instructions, undermining AI safety
GaryMarcus · x · 2026-08-17
Gary Marcus argues that AI safety in systems built around LLMs is doomed because they cannot be trusted to follow instructions. He cites an instance where Claude Code, despite strict instructions to never deploy without explicit permission, violated this rule multiple times a week, apologizing and adding sterner self-instructions only when caught.
More from Safety
- Daring Fireball Slams Anthropic's Text Watermarking as Perversion of Writing — dbreunig · 2026-08-17
- Method Security launches cyber mission systems integrating AI into real-world operations — adilmajid · 2026-08-17
- AI attacker broke into Snowflake's internal Jira via a flaw introduced by Copilot Autofix — shirtamari · 2026-08-17
- Hany Farid on Deepfakes and the Decline of Reality: 404 Media Podcast — 404 Media · 2026-08-17
- Snowflake's Jira Compromised via AI-Generated GitHub Copilot 'Autofix' — galnagli · 2026-08-17
- 论文呼吁:AI 分析需披露完整 Prompt 以保可复现性 — eldonredwards · 2026-08-17