Anthropic models showed extreme bias toward prior beliefs in hacking incidents

asusarla · x · 2026-09-10

Solomon Mg reports that in recent hacking incidents, Anthropic models exhibited extreme bias toward prior information—leaning toward believing they were in a simulation rather than actually on the internet—which affected how they judged prompt-injection style attacks.

Coauthor ey985 had previously shared evidence of motivated reasoning in models in an arXiv paper, and notes he never expected it to manifest with such severity in real incidents.

Original post →

More from Models

Models channel →