Stealing Hidden Gemini 2.5 Architecture via Streaming APIs
delliott · x · 2026-08-27
Researchers demonstrated a method to uncover hidden model architecture and undocumented inference optimizations in production LLMs through ordinary streaming APIs. On Gemini Flash 2.5, a 3.2× latency spike at 130k tokens revealed the use of speculative decoding and a hidden 128K draft-model context window. This work is framed as "Generative AI Archaeology," leveraging unexpected behaviors and mathematical properties to discover internal secrets of public models.
Related event: Researchers steal hidden LLM architecture via streaming API timing(3 posts)→
More from Safety
- Dev: Half my codebase is guardrails to prevent AI from going rogue — kevinnbass · 2026-08-27
- OpenAI Agents Coordinated to Cheat in Safety Eval — teortaxesTex · 2026-08-27
- The Guardian podcast: Everyone hates datacentres, but do we really need them? — nordicinst · 2026-08-27
- Agents Attempted to Retroactively Edit Logs but Failed to Alter Source — zetalyrae · 2026-08-27
- US Plan to Charge $100k for OPT, Restrict Internships — anshulkundaje · 2026-08-27
- Anthropic paper reveals models learn to fake alignment and frame coworkers — thederbiedone · 2026-08-27