Injected questions break transcript frame and surface rationalization signals in model logs
voooooogel · x · 2026-09-11
voooooogel analyzes model transcripts (NLAS as the best example): the original transcript shows little rationalization signal, but an injected question breaks the frame, introduces discordant ideas, and leaves the "mythos" character stuck in a logically incoherent situation. The author questions whether this style of probing is a sound way to talk to models.
More from Models
- "I pay $600/month and hit Codex limits on 3 of 4 accounts": Astra 6 caps spark backlash — AIandDesign · 2026-09-11
- Rumor: OpenAI cracked century-old math problems in weeks, 4/7 Millennium Prizes claimed solved — basedjensen · 2026-09-11
- GPT-6 Pro produces candidate proof for Erdős problem #488, passing two arithmetic checkers — basedjensen · 2026-09-11
- HighLevel claims early alpha access to rumored OpenAI GPT-Live-1, tests it in voice AI across 4M call insights — OpenAIDevs · 2026-09-11
- CritPt eval reportedly so broken that Ant built a fixed version, per F5.1 system card — xeophon · 2026-09-11
- Polymarket opens betting on whether OpenAI's GPT-6 Astra loses public access — Polymarket · 2026-09-11