Synthetic data can improve generalization when real data is lacking, but over-reliance can lead to distribution mismatches and degrade performance. In this paper, we introduce a learning theory framework for quantifying the trade-off between synthetic and real data. Our approach exploits the stability of the algorithm to derive generalized error bounds and characterize the optimal synthetic to real data ratio that minimizes the expected test error as a function of the Wasserstein distance between the real and synthetic distributions. We run the framework in a kernel ridge regression setting with mixed data, providing a detailed analysis that may be independently interesting. Our theory predicts the existence of an optimal ratio, leading to a U-shaped behavior of test errors for synthetic data ratios. We empirically validate this prediction on CIFAR-10 and clinical brain MRI datasets. Our theory extends to important scenarios of domain adaptation and shows that careful blending of synthetic target data and limited source data can reduce domain shifts and enhance generalization. Finally, we provide practical guidance for applying the results to both in-domain and out-of-domain scenarios.
- † University of Oxford
- ‡ UK Big Data Institute
