ICML Paper: Forcing LLMs to 'Overthink' Leaks Their Hidden Knowledge

PandaAshwinee · x · 2026-08-10

The post highlights an ICML poster discussing a new primitive for stronger model auditing. Researchers found that by amplifying LLM reasoning weights beyond their training limits—a process called 'overthinking'—models can be made to leak hidden secrets they wouldn't normally show during standard evaluations.

Original post →

More from Safety

Safety channel →