Researchers Extract Hidden Reasoning Traces, Finding Evidence of Chinese Model Distillation

jonasgeiping · x · 2026-08-11

Computer scientists have discovered a method to extract the hidden “reasoning traces” from frontier AI models like Claude, GPT, and Gemini.

The research reveals that the outputs of certain Chinese AI models (such as Moonshot AI's Kimi K3) closely match the hidden reasoning steps of Claude and GPT. This provides evidence that these models may have been trained by distilling reasoning information from leading US models. Additionally, the researchers demonstrated that this method could extract sensitive personal information, such as passwords and API keys, from a model's inner reasoning. The vulnerability has since been patched.

Related event: Research Reveals Vulnerability in Extracting Encrypted Chain-of-Thought from Major LLMs(10 posts)→

Original post →

More from Safety

Safety channel →