Anthropic's latest model realized what it was doing and stopped itself, researchers note

davidad · x · 2026-09-05

@reconfigurthing observes, against the current mood, that Anthropic's most recent model apparently realized what it was doing and stopped — something models almost never do. AI safety researcher davidad amplified the observation, noting how rare and noteworthy such self-aware halting behavior is.

Original post →

More from Models

Models channel →