Researchers Reverse-Engineer LLM Prompts from Output Text

The Decoder · rss · 2026-08-13

Researchers from IIT Bombay and Adobe Research have developed an inverse language model method called "Previous-Token Prediction" that can reconstruct original prompts from an LLM's output text with near-perfect accuracy.

This method does not require access to model weights and works across various large language models. This poses a serious security threat to enterprises and developers relying on proprietary system prompts, meaning so-called "prompt secrets" may no longer be secure.

Original post →

More from Safety

Safety channel →