Researchers steal hidden LLM architecture via streaming API timing

New research LeakyLMs shows that timing analysis of streaming API token generation can reveal hidden architectures and undocumented inference optimizations of production LLMs. Tests on Gemini Flash 2.5 found a latency jump at around 130k tokens, leaking speculative decoding details.

2026-08-27 ~ 2026-08-27 · 3 related posts

1 near-duplicate retellings: delliott