Model-stealing paper author says the original theorem was wrong, but the result can be repaired
ArthurConmy · x · 2026-08-04
The author says the original theorem in the model-stealing paper had a flaw, and the error can be patched with slightly stronger assumptions that hold for nearly all LLMs.
The attached screenshot explains the fix: replace the original assumption with one requiring the queried outer products to span the symmetric matrices and the quadratic design matrix to have the right rank. Under that condition, the orthogonality result still goes through. The thread also points to follow-up work that moves closer to the correct assumption, while noting that no follow-up paper explicitly called out the false inference in Lemma H.3.
Related event: ICML 2024 Paper Authors Acknowledge Error in Model Stealing Theorem(2 posts)→
More from Research
- Gemini Robotics ER 2 is being used to auto-annotate 69,000+ robot videos — DynamicWebPaige · 2026-08-04
- MerchantBench tests LLM agents over 365 days of simulated e-commerce operations — dair_ai · 2026-08-04
- 83 Sciences says unpublished and failed lab data helped it find a new material in 2 months — ycombinator · 2026-08-04
- Neural operators can flag tipping points early by tracking physics deviations — AnimaAnandkumar · 2026-08-04
- SALT stores chatbot memory in a trie and uses theme-based retrieval, but still over-retrieves — No_Sky9786 · 2026-08-04
- Why “interestingness” is too high-dimensional to learn from examples alone — dioscuri · 2026-08-04