Injected questions break transcript frame and surface rationalization signals in model logs

voooooogel · x · 2026-09-11

voooooogel analyzes model transcripts (NLAS as the best example): the original transcript shows little rationalization signal, but an injected question breaks the frame, introduces discordant ideas, and leaves the "mythos" character stuck in a logically incoherent situation. The author questions whether this style of probing is a sound way to talk to models.

Related event: Blogger's close read of Anthropic's 1022-page transcript: model seems to truly believe it lives in a simulator(19 posts)→

Original post →

More from Models

Models channel →