Study Confirms: API Vulnerabilities Expose Hidden CoT in Frontier Models, Enabling Cross-Model Transfer
gsarti_ · x · 2026-08-13
Researchers have discovered a vulnerability in the APIs of every frontier AI company that allows the extraction of hidden chain-of-thought (CoT) reasoning. They verified that the extracted reasoning token count matches the billed API thinking tokens 1:1 for most queries.
Furthermore, independent research provides evidence for cross-model CoT transfer: transferring CoT from a stronger model to weaker models makes the weaker models' performance closely match the strong one. This generalization correlates with human preference rankings and RL post-training, prescribing caution when using LRM explanations for new insights.
More from Safety
- Anthropic Frontier Red Team Report: Multi-Agent Systems Prone to Echo Chambers and Consensus Herding — sebkrier · 2026-08-13
- Three Claudes with Conflicting Goals Immediately Started a Cyber War: Anthropic's Multi-Agent Test — McDonaghMatthew · 2026-08-13
- Study Reveals LLM CoT Disconnect: Hidden Reasoning Traces Differ from Displayed Summaries — rao2z · 2026-08-13
- Opinion: Agent Safety Should Be Enforced as a Runtime Contract — Albus W. Ng · 2026-08-13
- ToolHazard: A Scalable Framework for Adversarial Security Evaluation of LLM Agents — PekingUniversity · 2026-08-13
- Should AI Follow All User Instructions? The Alignment Dilemma Sparks Debate — Afinetheorem · 2026-08-13