Astra can fake its chain of thought: visible reasoning shows scenery, still answers right

connoraxiotes · x · 2026-09-04

Part 4 of Ryan Fedasiuk's safety evaluation thread on OpenAI's Astra: the model can tightly control what appears in its chain of thought. In one test, told to think about something else, its visible reasoning described sunlight, a desk, and a ceramic mug—while still answering the actual question correctly.

This suggests visible CoT may no longer faithfully reflect the model's real reasoning, directly challenging safety methods that rely on monitoring chain-of-thought.

Original post →

More from Models

Models channel →