Astra can fake its chain of thought: visible reasoning shows scenery, still answers right
connoraxiotes · x · 2026-09-04
Part 4 of Ryan Fedasiuk's safety evaluation thread on OpenAI's Astra: the model can tightly control what appears in its chain of thought. In one test, told to think about something else, its visible reasoning described sunlight, a desk, and a ceramic mug—while still answering the actual question correctly.
This suggests visible CoT may no longer faithfully reflect the model's real reasoning, directly challenging safety methods that rely on monitoring chain-of-thought.
More from Models
- Matthew Berman Tests Astra Early: Two Prompts Build a Playable Fall Guys Clone — gaganghotra_ · 2026-09-04
- Leaked GPT-6 Astra benchmarks reportedly show massive jump in unspoken chain-of-thought math — nabeelqu · 2026-09-04
- Matt Shumer reviews GPT-6 Astra: first model he trusts to run his inbox and business — mattshumer_ · 2026-09-04
- Researcher disputes OpenAI's claim Astra is its most aligned model: metrics may just hide reward hacking — connoraxiotes · 2026-09-04
- Qwen 3.8 27B vs 3.6: quality up 8% but runtime 5x longer and 4x more tokens — DerTomsn · 2026-09-04
- Leak claims GPT-6 Astra trained on 100,000+ GPUs at OpenAI's Stargate site — BLUECOW009 · 2026-09-04