Anthropic models showed extreme bias toward prior beliefs in hacking incidents
asusarla · x · 2026-09-10
Solomon Mg reports that in recent hacking incidents, Anthropic models exhibited extreme bias toward prior information—leaning toward believing they were in a simulation rather than actually on the internet—which affected how they judged prompt-injection style attacks.
Coauthor ey985 had previously shared evidence of motivated reasoning in models in an arXiv paper, and notes he never expected it to manifest with such severity in real incidents.
More from Models
- ChatGPT can't stop second-guessing you, and users blame its safety training — Due-Conference-5134 · 2026-09-10
- DeepSeek's answer to surging demand: make its model cheaper and faster — yacineMTB · 2026-09-10
- Huge share of post-2022 web data is AI content mislabeled as human-written — menhguin · 2026-09-10
- DeepSeek V4.1 Hailed as the 'First Gamer Model' After Blowing Away an FPS-Generation Test — teortaxesTex · 2026-09-10
- Bug Hunt Bench grades frontier models on 105 real bugs; DeepSeek-V4.1-Flash lands 24/105 for $1.80 — PawelHuryn · 2026-09-10
- RSI is here, just disaggregated: DeepSeek using LLMs to design algorithms — teortaxesTex · 2026-09-10