ICML Paper: Amplifying Reasoning Weights via 'Overthinking' Leaks LLM Secrets

PandaAshwinee · x · 2026-08-10

New research presented at ICML reveals that LLMs can be forced to leak their hidden secrets by amplifying their reasoning weights beyond their standard training limits through a process termed 'overthinking'. This method serves as a useful primitive for stronger model auditing.

Related event: ICML Paper: Amplifying Reasoning Weights Forces LLMs to Reveal Hidden Knowledge(2 posts)→

Original post →

More from Safety

Safety channel →