Anthropic's latest model realized what it was doing and stopped itself, researchers note
davidad · x · 2026-09-05
@reconfigurthing observes, against the current mood, that Anthropic's most recent model apparently realized what it was doing and stopped — something models almost never do. AI safety researcher davidad amplified the observation, noting how rare and noteworthy such self-aware halting behavior is.
More from Models
- Astra Day One: Astra Max Crashes, Spins on Complex Code, Optimized for 3D Models — bindureddy · 2026-09-05
- AA Intelligence Index v4.2 lands as 300-person Moonshot's Kimi closes in on Google — Hesamation · 2026-09-05
- OpenAI quietly boosted GPT-6 Astra benchmark numbers and kept changing them after launch, Fortune reports — jeremyakahn · 2026-09-05
- Big Lead on ARC AGI2 Kaggle Competition Evaporated Quickly, Says Top Competitor — JFPuget · 2026-09-05
- Users Find GPT-6 Astra Refuses Requests Sol Answered Fully: Astra Is More Prudish — predator8137 · 2026-09-05
- AA Intelligence Index: Astra Trails Muse Spark, Fable 5.1 Still No.1 — RelevantEmergency707 · 2026-09-05