Two API Calls Exposed AI's Hidden Reasoning via New Attack

jbhuang0604 · x · 2026-08-24

Researchers demonstrated a novel attack against frontier models that extracts the model's encrypted hidden Chain of Thought using just two API calls. The attack works by replaying the encrypted reasoning traces through a weaker model to recover the hidden reasoning. The video showcases multiple examples of successfully extracted reasoning, revealing security vulnerabilities in how models output hidden thought processes.

Original post →

More from Safety

Safety channel →