Study Confirms: API Vulnerabilities Expose Hidden CoT in Frontier Models, Enabling Cross-Model Transfer

gsarti_ · x · 2026-08-13

Researchers have discovered a vulnerability in the APIs of every frontier AI company that allows the extraction of hidden chain-of-thought (CoT) reasoning. They verified that the extracted reasoning token count matches the billed API thinking tokens 1:1 for most queries.

Furthermore, independent research provides evidence for cross-model CoT transfer: transferring CoT from a stronger model to weaker models makes the weaker models' performance closely match the strong one. This generalization correlates with human preference rankings and RL post-training, prescribing caution when using LRM explanations for new insights.

Original post →

More from Safety

Safety channel →