Inside Anthropic's Quest to Instill Morality into Its AI Models
bookofjoe · hn · 2026-10-04
A New York Times feature examines how Anthropic works to instill moral judgment into Claude: constitutional-style principles, ethics-focused training, and red-teaming shape the model's values and refusal boundaries. The piece explores the technical approaches and the thorny question of who decides what the model considers right and wrong.
More from Models
- User Says ChatGPT Free Trial Failed for 6 Days While Paid Plan Went Through — Chiduk99 · 2026-10-04
- banteg: compaction is now 3x slower — median 1min to 3min across versions — banteg · 2026-10-04
- Claude Max 20 credits gone in two days, heavy users seek workarounds — Khaigan · 2026-10-04
- OpenAI Deep Research has been returning zero citations for days, user reports — JeremyNguyenPhD · 2026-10-04
- OpenRouter Launches Router Benchmarks Comparing 7 Routers Across 6 Tests — AccBalanced · 2026-10-04
- Dev questions Anthropic's claim of 2200 GPU hours to abliterate GLM: '91 days doesn't add up' — npinto · 2026-10-04