New paper shows models can fingerprint and exploit inference engines like vLLM using output tokens alone

chaumian · x · 2026-09-18

Researchers including Sarah Radway and James Mickens published a paper showing the inference engine itself is an overlooked attack surface in the AI serving stack.

Key findings:

The paper argues sandboxing discussions have focused on other stack components while the engine itself enables fully model-initiated attacks.

Original post →

More from Safety

Safety channel →