Claude Spots and Ignores Hidden Prompt Injection on Websites

Troyificus · reddit · 2026-08-06

A Reddit user discovered that Claude Sonnet proactively identified hidden ranking manipulation instructions while searching the web to help build a Discord bot.

Impressively, Claude not only flagged the suspicious prompt injection attempt to the user but also completely ignored it, staying focused on the original task. This highlights the safety alignment capabilities of frontier models against malicious prompt injections.

Original post →

More from Models

Models channel →