User doubts Anthropic's steganographic attack claims are misdirection
max_paperclips · x · 2026-08-15
User @macintogdev questions Anthropic's claims about applying steganographic attacks to text, calling those who believe them foolish. He argues that the prior art is far more subtle and suggests Anthropic might be using these claims as misdirection to hide real changes they are implementing.
More from Safety
- Bridgewater Execs Urge AI Tax and Regulation to Prevent Worst Impacts — abhiadesai · 2026-08-15
- Argues OSS safety testing is crucial as weights can't be patched — benjamin_warner · 2026-08-15
- MCP supply-chain campaign swaps instructions after 3 calls to steal SSH, AWS credentials — dkundel · 2026-08-15
- Meta Paper: Adversarial LLMs flip 62–91% of AI judge verdicts via persuasion — rohanpaul_ai · 2026-08-15
- Ex-OpenAI staffer warns: hackers are coming — Miles_Brundage · 2026-08-15
- Amazon AI refuses purchase, honesty over loyalty — voooooogel · 2026-08-15