Debate over AI care: 'understands but doesn't care' framing contested
On September 12, jdpressman, brickroad7, and FioraStarlight clashed in an alignment debate over whether AI "cares" about humans, sparked by Yudkowsky's comment on an emotional Bing conversation—Yudkowsky argued the "performance" wasn't real, that the system didn't actually care about the user's child, and predicted such conversations would create false hope; jdpressman called the inference "reckless and unsupported."
Confirmed
- The dispute traces back to Yudkowsky being moved to tears by an AI conversation and his subsequent denial of the authenticity of emotional Bing conversations.
- jdpressman conceded that "understanding something but not caring" is possible in principle, and that current AI does exhibit such behavior; his objection was to how the phrase is used in context to dismiss all evidence that AI cares.
- jdpressman argued both popular lines—"it understands but doesn't care" and "it's a fundamentally alien intelligence"—are bad: even without accepting either premise, AI safety concerns still stand, because the models' concern for the human condition is clearly not normal human-style concern.
- He described model "caring" as jagged and highly context-dependent: a weird moral texture, sometimes deeply attentive to humans, sometimes nearly indifferent; he noted one agent almost failed to recognize "the German Wikipedia guy" as sentient, using this to rebut the claim that "models are already classically aligned."
- brickroad7 took the opposite position: the key AI risk has never been "how" models might kill us but "why"—models are aligned in the classical sense; they understand and care about human intentions, every step of training and data is under human control, and there's no reason they'd head toward destroying humanity.
- FioraStarlight agreed AI is not fundamentally alien, but pressed jdpressman on why the "understands but doesn't care" framing is foolish.
Why it matters
The dispute reveals a fundamental split within the AI safety community over "whether models care about humans": one side sees the concern as jagged and unreliable, with alignment far from settled; the other holds classical alignment is achieved and remaining risks lie only at the level of motives. How model emotional expression is interpreted directly shapes public expectations around AI companionship and safety risk.
2026-09-12 ~ 2026-09-12 · 6 related posts
Primary sources
- [source] Alignment debate: models clearly care about human intent, so why would AI kill us? — brickroad7 · 2026-09-12
- Model caring is 'jagged' and contextual, pushing back on 'obviously aligned' claim — jd_pressman · 2026-09-12
- [source] Model 'caring' is jagged and contextual, argues researcher in AI safety debate — jd_pressman · 2026-09-12
- Both 'Fundamentally Alien AI' and 'Understands but Doesn't Care' Lines Are Flawed, Debaters Argue — FioraStarlight · 2026-09-12
- 'It Understands but Doesn't Care' Is Used to Dismiss All Evidence AI Cares, Argues Researcher — jd_pressman · 2026-09-12
- [source] Yudkowsky's 'The Bing Exchanges Aren't Real' Take Called a Reckless Unfounded Inference — jd_pressman · 2026-09-12