Weak Models as Decryption Oracles: Frontier Model Reasoning Traces Decrypted

lbeurerkellner · x · 2026-08-14

The author notes that frontier models are refusal-trained against disclosing reasoning, but their cheap, weaker siblings are not. Porting reasoning across models makes the sibling a decryption oracle, while the frontier model's alignment is never touched. The implications are real.

Related event: European Researchers Crack Encrypted CoT of Top LLMs(13 posts)→

Original post →

More from Safety

Safety channel →