Hidden LLM prompts can be reverse-engineered from outputs alone
mhmazur · x · 2026-08-05
Researchers have introduced PTP (Previous-Token Prediction), an attack method capable of reconstructing a language model's hidden system prompt solely by observing its generated text.
By exploiting the statistical trail left by next-token prediction, the method trains an inverse model to deduce preceding tokens for a given response. This implies that the hidden rules of customer support bots could be extracted just by reading their public replies. However, the attack's effectiveness drops significantly if the target model's output is highly generic.
More from Safety
- Dean Ball Rebuts Doomers: Moderate Prudence Will Make Transformative AI Worth It — AaronBergman18 · 2026-08-05
- Dev Questions Legality of AI Apps Getting Full Customer Data API Access — jonathan_wilke · 2026-08-05
- Clarification: Anthropic Proactively Withdrew from China, Not Blocked — teortaxesTex · 2026-08-05
- US Federal AI Budget Hits $90.7B, with 98.9% Flowing to the DoD — ChrisUniverse · 2026-08-05
- Founder Returns to Startup to Build Access Control for Autonomous Agents — gabriel1 · 2026-08-05
- Stanford HAI Scholars Warn World Models Pose Physical Risks and Demand New Governance — StanfordHAI · 2026-08-05