Researchers steal hidden LLM architecture via streaming API timing
New research LeakyLMs shows that timing analysis of streaming API token generation can reveal hidden architectures and undocumented inference optimizations of production LLMs. Tests on Gemini Flash 2.5 found a latency jump at around 130k tokens, leaking speculative decoding details.
2026-08-27 ~ 2026-08-27 · 3 related posts
- LeakyLMs: Stealing Architecture and Inference Optimizations via Timing — niloofar_mire · 2026-08-27
- Stealing Gemini 2.5's Hidden Architecture via Streaming API Side-Channel — alexisjross · 2026-08-27
1 near-duplicate retellings: delliott