Google rebuilds federated learning on TEEs: 3x smaller privacy budget for Gboard, trained in 3 weeks

Toward provably private learning from federated data

Katharine Daly, Yu Xiao, Zachary Garrett, Brett McLarnon, Jianpeng Hou, Arun Ganesh, Yanxiang Zhang, Noriyuki Takahashi, Haicheng Sun, Yuanbo Zhang, Timon Van Overveldt, Daniel Ramage

cs.CR

2026-09-26

Devices upload encrypted data; training runs in server-side TEEs. Gboard models now train in 3 weeks, not 2 months, at a 3x smaller externally verifiable privacy budget.

What problem this solves

Federated learning has backed on-device features at Google and Apple since 2016: next-word prediction, emoji suggestion, speech recognition. The standard design has devices compute gradients locally and a server aggregate them. That design carries three long-standing costs. Model size is capped by what a phone can train. Synchronizing fleets of devices with unpredictable availability is operationally expensive. And the privacy story runs on trust: users and outside auditors cannot see the server-side code, so they cannot check that DP noise was applied as claimed.

Google redefined FL in 2024 around four privacy principles: data minimization, data anonymization, transparency and control, and verifiability and auditability. This paper describes the system built to that standard, already running in production.

Method

The core swap is singular: devices no longer upload gradients protected by secure aggregation. They upload raw data encrypted under keys from a TEE-hosted KMS, and training moves entirely into server-side TEEs built on AMD SEV-SNP and Intel TDX.

Uploading raw data buys freedom. The DP training loop runs after all uploads are in, so participation no longer depends on device availability. Every device can participate an equal number of times, the minimum round separation (minSep) reaches its maximum for the collected data, and the BLT noise matrix (banded lower-triangular correlated noise) and multiplier follow from the target zCDP. Less noise for the same privacy budget.

Results

Device coverage, Japanese model, 3000 rounds, cohort 6500:

MetricOld systemNew TEE system
Devices used8.5M of 35.5M available over 38 days (23.9%)17.8M uploads in about 6 days, all used
Why devices missed20.5M never got a task; 6.5M interruptedMostly no eligible data

Privacy-utility, English model, both MF-DP-FTRL, cohort 6500:

MetricOld systemNew TEE system
Noise multiplier at zCDP=0.232, T=50009.545.16
maxP / minSep at T=50008 / 5613 / 1822
Setup8616 rounds, 85 days, fixed multiplier 7.385000 rounds, parameters set at runtime from 11.8M uploads

Live A/B with 3.5M devices per arm: the best TEE arm trained at zCDP=0.215, one third of the production model's 0.641 adjusted to the same mechanism, with neutral Words Per Minute and Words Modified Ratio. Training took 3 weeks instead of 2 months.

Why it matters

The paper presents this as the first production FL system with externally verifiable central DP guarantees, and the mechanism supports the claim. DP statements used to be "trust us"; now the program and binary hashes sit in a public log anyone can compare against what actually runs. The engineering gains are just as real: training is untethered from device compute and availability, participation across 17.8M uploads is unbiased, and the cycle compresses from two months to three weeks. For teams building privacy-preserving ML, this is a working alternative to SecAgg-style cryptography: trust moves from hard math to hardware attestation, and it buys verifiability plus a better privacy-utility curve.

Limitations

The authors' own list: current-generation TEEs have known security weaknesses; proprietary per-user processing does not break the verifiable DP guarantee but remains exposed to side channels; one hyperparameter configuration per model was tested, and signal-to-noise ratio drops sharply at maxP transitions; the pipeline operator can still see which (encrypted) uploads feed each round, with sampling-based amplification left to future work; the TTL forces a tradeoff between waiting for uploads and leaving time to train. Scale stops at 10M-parameter models so far; larger ones need GPU-backed worker TEEs.

Two caveats from a close read: the verifiability ultimately rests on AMD and Intel hardware promises, and side channels are excluded from the guarantee rather than solved. The A/B reports neutral typing metrics only, with no direct accuracy comparison, so the abstract's "better accuracy" claim has no matching number in the A/B section.

Terms

Source

What people are saying

Related papers

All paper explainers