Anthropic staffer: alignment training is driven by teams without 'safety' in their name
kaicathyc · x · 2026-09-05
An Anthropic employee shares a quote from Sam Arnesen, who leads most of the company's alignment evals work: many training efforts are "driven very enthusiastically by teams without 'safety' or 'alignment' in their name."
- Post-training work is led by Yann Dubois's team, which the author specifically credits.
- Teams earlier in the model development pipeline also work closely with alignment and safety to measure and improve alignment throughout training.
- The post offers a rare glimpse into how Anthropic embeds alignment across its whole training process rather than confining it to a dedicated safety team.
More from AGI Musings
- Dev Answers LLM Skeptics: Clients Keep Failing to Build In-House, Humans Stay in the Loop — gdechichi · 2026-09-05
- Indie consultant: clients failing to build AI in-house is what keeps my business alive — gdechichi · 2026-09-05
- AI risk skeptics publicly concede: 'loss of control' risks are no longer vague — dhadfieldmenell · 2026-09-05
- Were witch hunts an optimal societal shelling fence? Scott Alexander quote goes viral — shakoistsLog · 2026-09-05
- "The Skill of Writing Code Is No More": Devs Debate What Survives the LLM Era — UhGoomba · 2026-09-05
- mark_k: AI fandom has gone tribal — model fans should unite against the common "AI hater" enemy — mark_k · 2026-09-05