Singular-limit analysis of gradient descent with noise injection

Anna Shalova, André Schlichting and Mark A. Peletier

Accepted in Journal of Machine Learning Research

Abstract

We study the limiting dynamics of a large class of noisy gradient descent systems in the overparameterized regime. In this regime the set of global minimizers of the loss is large, and when initialized in a neighbourhood of this zero-loss set a noisy gradient descent algorithm slowly evolves along this set. In some cases this slow evolution has been related to better generalisation properties. We characterize this evolution for the broad class of noisy gradient descent systems in the limit of small step size. Our results show that the structure of the noise affects not just the form of the limiting process, but also the time scale at which the evolution takes place. We apply the theory to Dropout, label noise and classical SGD (minibatching) noise, and show that these evolve on different two time scales. Classical SGD even yields a trivial evolution on both time scales, implying that additional noise is required for regularization. The results are inspired by the training of neural networks, but the theorems apply to noisy gradient descent of any loss that has a non-trivial zero-loss set.

Noisy gradient descent may continue to move after reaching the zero-loss set Γ. The left-hand panel shows the level curves of a function L:^2→[0,∞), with the zero-level set Γ marked in red. The middle panel shows a gradient-descent evolution, starting at the top, and converging to Γ. The right-hand panel shows an evolution of the noisy gradient descent with L̂(w,η) := L(w+η).
Noisy gradient descent may continue to move after reaching the zero-loss set Γ\Gamma . The left-hand panel shows the level curves of a function L:R2[0,)L:\R^2\to[0,\infty) , with the zero-level set Γ\Gamma marked in red. The middle panel shows a gradient-descent evolution, starting at the top, and converging to Γ\Gamma . The right-hand panel shows an evolution of the noisy gradient descent with L^(w,η):=L(w+η)\hat L(w,\eta) := L(w+\eta) .

Publication history

Preprint
2024-04-18
Accepted
2026-07-17
Published
2026-07-17

Topics: machine learning, gradient flows, variational convergence