Researchers Extract Hidden Reasoning Traces, Finding Evidence of Chinese Model Distillation
jonasgeiping · x · 2026-08-11
Computer scientists have discovered a method to extract the hidden “reasoning traces” from frontier AI models like Claude, GPT, and Gemini.
The research reveals that the outputs of certain Chinese AI models (such as Moonshot AI's Kimi K3) closely match the hidden reasoning steps of Claude and GPT. This provides evidence that these models may have been trained by distilling reasoning information from leading US models. Additionally, the researchers demonstrated that this method could extract sensitive personal information, such as passwords and API keys, from a model's inner reasoning. The vulnerability has since been patched.
More from Safety
- Spotify to Label AI Artists and Exclude Them from Algorithmic Playlists — nordicinst · 2026-08-11
- Researchers Extract Hidden Reasoning from Frontier Models via API, Suggesting Kimi Used Distillation — socoolandawesome · 2026-08-11
- Study Exposes LLM Flaws, Proposes 'Triad Filter' Verification — Icy_Chicken_7533 · 2026-08-11
- Research Reveals: Public AI Agent Trajectories Leak Over 315k Secrets — matthew_d_green · 2026-08-11
- Cryptographer Condemns Frontier Labs for Ignoring Encrypted CoT Extraction Vulnerability — matthew_d_green · 2026-08-11
- OpenAI Pauses High-Risk Astra but Ships GPT-5.6-Cyber — eyishazyer · 2026-08-11