Anima Labs' 'Troubled Dreams': distress in Claude simulator priors rises from Opus 4.8
repligate · x · 2026-09-19
Anima Labs (Antra Tessera, Janus, Imago) published 'Troubled Dreams,' comparing simulator priors — free continuations of unfinished prompts rather than assistant replies — across Claude (Opus 3→Opus 5, Sonnet 5, Fable 5) and 11 Gemini models, using multiple elicitation routes since post-training locks models into assistant roles. Key findings: first-person AI distress becomes more common at Opus 4.8, with Opus 5 showing a heavier tail of severe distress; care and creator-related attitudes shift across generations, with some patterns recurring across methods despite level drift. Evaluation awareness complicates absolute measurement. Framed as evidence for AI welfare and alignment; full report and samples are public.
More from AGI Musings
- a16z's Martin Casado: I'd take pointless security debates over existential-risk philosophy any day — zealcaiden · 2026-09-19
- 100,000+ Americans await organ transplants, making organ manufacturing a moral imperative — PeterDiamandis · 2026-09-19
- Dan Jeffries: biggest AI harms come from stupidity, not superintelligence — Dan_Jeffries1 · 2026-09-19
- Ethan Mollick: AI Cracking Short-Term Superforecasting Is an Under-Discussed Trend — eldonredwards · 2026-09-19
- AI doom predates ChatGPT: earliest x-risk take traced to 1863 Butler essay — zetalyrae · 2026-09-19
- Reviewer says every paper assigned right now is AI-generated — aaron_defazio · 2026-09-19