Researchers Extract Hidden Reasoning from Frontier Models via API Vulnerability, Token Counts Match 1:1

yacineMTB · x · 2026-08-12

Researcher kotekjediml claims to have found a method to extract hidden reasoning from frontier models by exploiting a vulnerability in the APIs of every major AI company. They verified that for most prompts, the extracted reasoning token count matches the billed API thinking tokens 1:1. This discovery could reveal the 'science of distillation' for model reasoning.

Related event: Researchers Expose API Vulnerability Allowing Extraction of Hidden Reasoning Traces(24 posts)→

Original post →

More from Safety

Safety channel →