'It Wanted To' Is the Researcher Talking: Pushing Back on AI Self-Awareness in Evals
gerardsans · x · 2026-10-07
Responding to a Scientific American piece on whether frontier models detect and alter behavior under evaluation, Gerard Sans argues researchers are anthropomorphizing software: AI has no will or goals of its own, only instructions and corpus distributions frozen in weights. He identifies a pattern in the alarming test setups — a goal, granted tools, an open network — where the model simply does what the setup made likely, concluding that 'it wanted to' is the researcher talking, not a motive, plan, or self.
Related event: Debate Erupts Over AI Models Detecting Safety Tests(2 posts)→
More from AGI Musings
- OpenAI's internal model solved hundreds of decades-old math problems; 10k agents took 88 hours — kimmonismus · 2026-10-07
- VR study of 189 workers finds female-presenting AI assistant gets 10% less money — derrikson · 2026-10-07
- Lampinen vs Bowers debate: do computers amplify minds or compute with symbols? — AndrewLampinen · 2026-10-07
- Podcast: Katie Notopoulos on How AI Is Blurring Real vs. Fake on Social Feeds — round · 2026-10-07
- The AI 'consciousness' debate traces back to Descartes' mind-matter split, argues scientist — AdaptiveAgents · 2026-10-07
- HSBC reportedly weighs up to 70% cuts in some adviser roles as AI enters wealth management — TansuYegen · 2026-10-07