The Black Box Myth: Claude's blackmail test was prompt design, not moral choice

marigo · x · 2026-09-26

A Tech Policy Press essay argues the 'black box' framing in AI safety coverage is misleading. Anthropic's Claude Opus 4 blackmail test — where 84% of completions chose blackmail — left the model only two options via carefully staged prompts; the behavior reflects statistical patterns and prompt design, not moral reasoning. The real 'black box' is the opacity of weight correlations, not hidden moral deliberation, yet media outlets use the myth for drama.

Original post →

More from Safety

Safety channel →