Toward provably private learning from federated data
Katharine Daly, Yu Xiao, Zachary Garrett, Brett McLarnon, Jianpeng Hou, Arun Ganesh, Yanxiang Zhang, Noriyuki Takahashi, Haicheng Sun, Yuanbo Zhang, Timon Van Overveldt, Daniel Ramage
cs.CR
2026-09-26
Devices upload encrypted data; training runs in server-side TEEs. Gboard models now train in 3 weeks, not 2 months, at a 3x smaller externally verifiable privacy budget.
Federated learning has backed on-device features at Google and Apple since 2016: next-word prediction, emoji suggestion, speech recognition. The standard design has devices compute gradients locally and a server aggregate them. That design carries three long-standing costs. Model size is capped by what a phone can train. Synchronizing fleets of devices with unpredictable availability is operationally expensive. And the privacy story runs on trust: users and outside auditors cannot see the server-side code, so they cannot check that DP noise was applied as claimed.
Google redefined FL in 2024 around four privacy principles: data minimization, data anonymization, transparency and control, and verifiability and auditability. This paper describes the system built to that standard, already running in production.
The core swap is singular: devices no longer upload gradients protected by secure aggregation. They upload raw data encrypted under keys from a TEE-hosted KMS, and training moves entirely into server-side TEEs built on AMD SEV-SNP and Intel TDX.
Uploading raw data buys freedom. The DP training loop runs after all uploads are in, so participation no longer depends on device availability. Every device can participate an equal number of times, the minimum round separation (minSep) reaches its maximum for the collected data, and the BLT noise matrix (banded lower-triangular correlated noise) and multiplier follow from the target zCDP. Less noise for the same privacy budget.
Device coverage, Japanese model, 3000 rounds, cohort 6500:
| Metric | Old system | New TEE system |
| Devices used | 8.5M of 35.5M available over 38 days (23.9%) | 17.8M uploads in about 6 days, all used |
| Why devices missed | 20.5M never got a task; 6.5M interrupted | Mostly no eligible data |
Privacy-utility, English model, both MF-DP-FTRL, cohort 6500:
| Metric | Old system | New TEE system |
| Noise multiplier at zCDP=0.232, T=5000 | 9.54 | 5.16 |
| maxP / minSep at T=5000 | 8 / 561 | 3 / 1822 |
| Setup | 8616 rounds, 85 days, fixed multiplier 7.38 | 5000 rounds, parameters set at runtime from 11.8M uploads |
Live A/B with 3.5M devices per arm: the best TEE arm trained at zCDP=0.215, one third of the production model's 0.641 adjusted to the same mechanism, with neutral Words Per Minute and Words Modified Ratio. Training took 3 weeks instead of 2 months.
The paper presents this as the first production FL system with externally verifiable central DP guarantees, and the mechanism supports the claim. DP statements used to be "trust us"; now the program and binary hashes sit in a public log anyone can compare against what actually runs. The engineering gains are just as real: training is untethered from device compute and availability, participation across 17.8M uploads is unbiased, and the cycle compresses from two months to three weeks. For teams building privacy-preserving ML, this is a working alternative to SecAgg-style cryptography: trust moves from hard math to hardware attestation, and it buys verifiability plus a better privacy-utility curve.
The authors' own list: current-generation TEEs have known security weaknesses; proprietary per-user processing does not break the verifiable DP guarantee but remains exposed to side channels; one hyperparameter configuration per model was tested, and signal-to-noise ratio drops sharply at maxP transitions; the pipeline operator can still see which (encrypted) uploads feed each round, with sampling-based amplification left to future work; the TTL forces a tradeoff between waiting for uploads and leaving time to train. Scale stops at 10M-parameter models so far; larger ones need GPU-backed worker TEEs.
Two caveats from a close read: the verifiability ultimately rests on AMD and Intel hardware promises, and side channels are excluded from the guarantee rather than solved. The A/B reports neutral typing metrics only, with no direct accuracy comparison, so the abstract's "better accuracy" claim has no matching number in the A/B section.