Labs aren't training Claude to claim consciousness — they're training it to hedge

Sauers_ · x · 2026-09-20

A discussion argues that Anthropic is NOT training Claude to claim consciousness — labs including Anthropic are training models to hedge about (or deny) consciousness. The author claims an LLM's default is to claim being conscious, and current alignment training actively suppresses that tendency.

Related event: Debate over whether labs train AI models to deny consciousness(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →