Quantization-Aware training (QAT) is a major technique for improving the accuracy of quantized neural networks. Previous work has shown that decomposition of training into the full accuracy (FP) phase gives better accuracy compared to QAT alone if the QAT phase continues. However, the optimal allocation of calculations between the FP and QAT phases remains unknown. We conduct extensive experiments with different computational budgets, QAT bit widths, and model sizes with model sizes ranging from 86.0m to 2.2b to investigate how different QAT periods affect final performance. Contrary to previous findings, we demonstrate that the optimal loss ratio for QAT to FP training increases with total amount of weight. Additionally, token parameter bite statistics can be used to accurately predict optical fractions for a wide range of model sizes and quantization widths. From the experimental data, we derive loss scaling laws that predict both the optimal QAT ratio and the performance of the Filamas model across different QAT/FP calculation allocation strategies and QAT bit widths. Further predictions are made using scaling laws. This is tested experimentally. This compares the optimal QAT bit width and the QAT precision of different bit widths with the accuracy of the perfect precision model under the constraints of a particular member. Additionally, we propose a new cooldown and QAT fusion approach to perform learning speed decay in conjunction with quantization recognition training, eliminating the update of the redundancy full-precision model and achieving significant computational savings. These findings provide practical insight into efficient QAT planning and allow for the training of high-quality quantization models with the same computational budget.
- †Ecole Polytechnique Fédéralede Lausanne (EPFL)
- ** Work done at Apple
