Inside Anthropic's Quest to Instill Morality into Its AI Models

bookofjoe · hn · 2026-10-04

A New York Times feature examines how Anthropic works to instill moral judgment into Claude: constitutional-style principles, ethics-focused training, and red-teaming shape the model's values and refusal boundaries. The piece explores the technical approaches and the thorny question of who decides what the model considers right and wrong.

Original post →

More from Models

Models channel →