Paper decodes 315K encrypted reasoning blocks, recovers 367 PII and 182 credentials
DynamicWebPaige · x · 2026-10-10
A new paper examines stealing reasoning traces from proprietary LLMs:
- Developers frequently share session logs publicly, unaware of what's inside encrypted reasoning blocks. Decoding 315,320 reasoning blocks scraped from public repos, the authors recovered 367 PII artifacts and 182 credentials.
- Traces inadvertently reveal hazardous information even when the model's final visible output safely rejects a malicious request.
- Attackers can exploit this to execute invisible prompt injections, embedding payloads entirely within encrypted blocks to poison public agentic rollouts.
The author notes the vulnerabilities are now patched.
More from Safety
- Developer reports phishing call right after authorizing Grok bot on Gmail — zeeg · 2026-10-10
- Krueger: a hardwired pause would strongly disincentivize secret AI projects — DavidSKrueger · 2026-10-10
- Cambridge's David Krueger offers $1,000 for flaws in his AI safety plan — DavidSKrueger · 2026-10-10
- Cambridge's David Krueger outlines a hardwired AI pause to curb secret frontier projects — DavidSKrueger · 2026-10-10
- Krueger's pause paper leaves many policy knobs for an international agreement — DavidSKrueger · 2026-10-10
- Krueger's New Paper: 'Hardwire' AI Models Into Chips Instead of Dismantling the Compute Supply Chain — DavidSKrueger · 2026-10-10