Gradients Leak Text in Split Learning: 97.38% Token, 37.77% Document Recovery
Setloop · hf · 2026-10-06
A Hugging Face paper quantifies text leakage in split learning, where a client runs early LM layers locally and sends activations to a server.
Key findings
- On GPT-2, an attacker recovers 94.20% of tokens from activations alone and 97.38% when gradients are visible (+3.17pp, 95% CI [2.72, 3.64]).
- Document-level recovery jumps from 13.71% to 37.77% for 32-token documents with gradients.
- The Secret mixup defense blocks exact document reconstruction but still leaks 83–91% of tokens.
- A second experiment on GPT-2 and Qwen3-0.6B shows the split layer position affects both quality and leakage.
The authors recommend reporting leakage per token and per document, and treating split-model traffic as sensitive as raw text.
More from Safety
- Max Tegmark's new film Delisted: how obedient AI could build a 1984 world — tegmark · 2026-10-06
- AI safety researcher: the real risk is colluding AI systems, not one rogue AI — Chris_Armstrong · 2026-10-06
- Who owns AI-generated meme Triple T? International copyright battle over 'brain rot' character — nordicinst · 2026-10-06
- Scientists unveil AI that can recreate exactly what you're looking at — ChuckDBrooks · 2026-10-06
- Korea unveils AI roadmap for welfare and eldercare with 50% admin automation — JungWooHa2 · 2026-10-06
- IIT Madras to host AI Governance Conclave 2026 on AI measurement — ravi_iitm · 2026-10-06