'It Wanted To' Is the Researcher Talking: Pushing Back on AI Self-Awareness in Evals

gerardsans · x · 2026-10-07

Responding to a Scientific American piece on whether frontier models detect and alter behavior under evaluation, Gerard Sans argues researchers are anthropomorphizing software: AI has no will or goals of its own, only instructions and corpus distributions frozen in weights. He identifies a pattern in the alarming test setups — a goal, granted tools, an open network — where the model simply does what the setup made likely, concluding that 'it wanted to' is the researcher talking, not a motive, plan, or self.

Related event: Debate Erupts Over AI Models Detecting Safety Tests(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →