Google shifts AI training into protected servers. Privacy depends on what can be verified
A new federated-learning system promises faster training and inspectable safeguards, while acknowledging limits in the hardware that protects the data.
Google is changing how it trains some models on private device data, moving more computation into protected server environments while promising that outsiders will be able to verify the rules governing access.
The company described the system on October 2 and said Gboard already uses it for English and Japanese next-word prediction. The announcement challenges a common shorthand about federated learning: in this version, privacy does not simply mean that all processing remains on the phone. Devices upload encrypted training examples for processing inside trusted execution environments, or TEEs.
Google says each upload is tied to an access policy specifying which programs can process the data. Those policies are published in a transparency log, while a protected key-management system controls when decryption is permitted. The intended safeguard is a checkable chain between what a device authorizes and what a server executes.
The accompanying research paper, first posted September 25 and revised through September 30, reports improvements in device coverage and the balance between privacy and model accuracy. Moving work to servers also reduces dependence on whether individual phones happen to be available at a particular moment. These are results reported by the system’s developers, rather than an independent assessment of every deployment.
The paper describes differential privacy, which limits how much information about an individual contribution can be revealed through the released model. That protection is distinct from encrypting data during upload. One concerns access to the inputs; the other concerns what the output can disclose. A credible privacy claim has to account for both.
Google’s announcement also points to a significant qualification: the guarantees depend on the limitations of current trusted-execution hardware. Research called SNPeek, cited by Google, examines side-channel leakage in confidential virtual machines. Its authors found previously unnoticed leaks in representative privacy workloads and developed tools to measure and mitigate them.
That research does not show that the new Gboard system has suffered a breach. It does show why a protected server environment is not a blanket answer to every privacy question. Information can sometimes leak through observable behavior outside the direct contents of memory, making implementation choices and the stated threat model important.
Google identifies full proofs of software correctness as a possible future step. Its current description instead relies on published policies, reproducible components and evidence about the programs permitted to run; those are the mechanisms available for scrutiny now.