Google presents VaultGemma, an AI model that protects sensitive data without compromising performance. The 1 billion parameter model can be used as open source using privacy differences.
Google Research and Google Deepmind are behind Vaultgemma, a language model that solves traditional AI privacy issues. The model is built on Google's Gemma Architecture, indicating that differences in privacy do not necessarily mean poor performance.
Differential privacy works by adding control noise to the dataset. This makes it impossible to obtain certain information while maintaining overall ease of use. Vaultgemma was built from scratch and trained with a discriminatory privacy framework to ensure that sensitive data is not remembered or leaked.
New scaling methods break through old limits
Traditional scaling methods for AI models do not apply when privacy differences are applied. Therefore, Google has developed a new “DP scaling law” that takes into account additional noise and larger batch sizes. This breakthrough allows for the development of larger and more powerful personal language models.
The team adapted the training protocol to counter the instability caused by the addition of noise. Private models require batch sizes with millions of examples to train stably. Google has found a way to reduce these computational costs without compromising privacy guarantees.
Performance comparable to public models
For evaluations on benchmarks such as MMLU and Big-Bench, VaultGemma works on par with non-private Gemma models with the same number of parameters. This is surprising as the previous differential private models have always deteriorated significantly.
VaultGemma uses a 26-layer and multi-query attention-grabbing decoder-only trans architecture. The length of the sequence is limited to 1,024 tokens to keep the intensive computational requirements of private training manageable.
Open source for wider adoption
Google is making Vaultgemma completely open source by hugging Face and Kaggle. This contrasts with proprietary models such as the Gemini Pro. The new scaling method must be applicable to much larger private models. Google envisions collaboration with healthcare providers using VaultGemma, which analyses sensitive patient data without risk of privacy.
By refusing to disclose training data, the model also reduces the risk of misinformation and bias amplification, the researchers say.
Tip: Google mainly places Gemini behind the paywall
