A model-stealing paper’s proof needs a stronger spanning assumption, the author says

ArthurConmy · x · 2026-08-04

This follow-up thread links the paper and the discussion chat, then clarifies that the author’s earlier claim had an error in the proof, though the main takeaway still stands.

The quoted paper is Stealing Part of a Production Language Model, which shows a black-box attack that can recover precise information from production LLM APIs, including embedding projection layers, hidden dimension sizes, and, for some models, the full projection matrix at low cost. The thread’s correction says the proof needs a stronger spanning assumption and notes related papers that move toward the right condition, but do not directly identify the flawed lemma.

Original post →

More from Research

Research channel →