Stealing Hidden Gemini 2.5 Architecture via Streaming APIs

delliott · x · 2026-08-27

Researchers demonstrated a method to uncover hidden model architecture and undocumented inference optimizations in production LLMs through ordinary streaming APIs. On Gemini Flash 2.5, a 3.2× latency spike at 130k tokens revealed the use of speculative decoding and a hidden 128K draft-model context window. This work is framed as "Generative AI Archaeology," leveraging unexpected behaviors and mathematical properties to discover internal secrets of public models.

Related event: Researchers steal hidden LLM architecture via streaming API timing(3 posts)→

Original post →

More from Safety

Safety channel →