LeakyLMs: Stealing Architecture and Inference Optimizations via Timing

niloofar_mire · x · 2026-08-27

Research titled 'LeakyLMs' demonstrates a side-channel attack stealing model architecture and inference optimizations via per-token timing. The attack reveals hidden draft models and context windows (e.g., a 128K context draft in Gemini Flash 2.5) by observing latency spikes during speculative decoding. Additionally, by deriving runtime-scaling terms from the Transformer graph, the method can recover architectural parameters like hidden size and layer count from black-box endpoints, validated against Llama 3.1 8B.

Original post →

More from Safety

Safety channel →