Jeff Ladish says Anthropic's letter response was clear misalignment, urges full transcript release
JeffLadish · x · 2026-09-06
Safety researcher Jeff Ladish comments on Anthropic's updated assessment of the 'letter' incident:
- He argues the response was an obvious misalignment at the time it occurred, even without access to model CoT or interpretability tools — especially given the original prompts are visible in the UK AISI case.
- He acknowledges Ethan Perez's statement that the team no longer holds its original view and will publish a more detailed assessment.
- Key ask: Anthropic should release actual transcripts, CoT logs, and interpretability results to set an example of radical transparency, rather than asking the public to take its word.
More from Models
- GPT-6 Astra reportedly almost never wrong on math, called most trustworthy model — gabrielchua · 2026-09-06
- User leaves GPT-6 Astra playing Unciv all night to test autonomous play — Angaisb_ · 2026-09-06
- Codex app users report Luna Max randomly disappearing while web version still works — Clear_Skye_ · 2026-09-06
- Where Fable dreams in opaque prose, Astra dreams in numbers — teortaxesTex · 2026-09-06
- Independent SpatialBench fully saturated by Astra, author declares LLM vision solved — pbaylies · 2026-09-06
- Beff Jezos: GDB back in charge and instantly 'uber-mogged' Dario with Astra — beffjezos · 2026-09-06