Hidden LLM prompts can be reverse-engineered from outputs alone

mhmazur · x · 2026-08-05

Researchers have introduced PTP (Previous-Token Prediction), an attack method capable of reconstructing a language model's hidden system prompt solely by observing its generated text.

By exploiting the statistical trail left by next-token prediction, the method trains an inverse model to deduce preceding tokens for a given response. This implies that the hidden rules of customer support bots could be extracted just by reading their public replies. However, the attack's effectiveness drops significantly if the target model's output is highly generic.

Original post →

More from Safety

Safety channel →