Testing whether Claude models notice when you swap models mid-conversation — Fable didn't refuse
davidmanheim · x · 2026-09-23
David Manheim probes model self-awareness: asking Sonnet about a likely-refused-by-Fable topic (synthesis of dithiophosphates, barely related to nerve agents), then switching to Fable — which didn't refuse. The core question raised: can models "notice" that preceding text isn't something they'd write and act on it unprompted.
Related event: Cross-Model Experiment: Sonnet Answers Sensitive Topic That Fable Refuses(2 posts)→
More from Models
- Testing the Jeb chatbot: inconsistently biased, not neutral — calibrate it like any classifier — PawarBI · 2026-09-24
- Pokemon benchmark Paradigm 3: Astra generalizes to scrambled maps and fan-made games while rivals memorize — gleech · 2026-09-24
- AI Completes Fan-Made Pokemon Brown in 10K Steps: Real Generalization or Whack-a-Mole? — gleech · 2026-09-24
- Next-gen model names surface: Opus 5.5, Fable 5.1, GPT-6 Astra — labs said to be ~2 months ahead internally — haider1 · 2026-09-24
- AI detector debate: economist argues Pangram is the only reliable tool, cites 0 FPR finding — paulnovosad · 2026-09-24
- MiniMax H3 on Spectrum Runs Each Step Twice, Killing Performance — Glittering-Cold-2981 · 2026-09-24