- Featured in Physics
- Open Access
Analytic theory of dropout regularization
Phys. Rev. E 112, 045301 – Published 1 October, 2025
DOI: https://doi.org/10.1103/jmdx-x3gr
Abstract
Dropout is a regularization technique widely used in training artificial neural networks to mitigate overfitting. It consists of dynamically deactivating subsets of the network during training to promote more robust representations. Despite its widespread adoption, dropout rates are often selected heuristically, and theoretical explanations of its success remain sparse. Here we analytically study dropout in two-layer neural networks trained with online stochastic gradient descent. In the high-dimensional limit, we derive a set of ordinary differential equations that fully characterize the evolution of the network during training and capture the effects of dropout. We obtain a number of exact results describing the generalization error and the optimal dropout probability at short, intermediate, and long training times. Our analysis shows that dropout reduces detrimental correlations between hidden nodes and mitigates the impact of label noise and that the optimal dropout probability increases with the level of noise in the data. Our results are validated by extensive numerical simulations.
Physics Subject Headings (PhySH)
Collections
This article appears in the following collection:

Statistical Physics Meets Machine Learning - Machine Learning Meets Statistical Physics
The Editors of Physical Review E are pleased to present the Collection on Statistical Physics Meets Machine Learning - Machine Learning Meets Statistical Physics, highlighting research at the intersection of machine learning and statistical physics, on the occasion of the two Statistical Physics Meets Machine Learning and the two Machine Learning Meets Statistical Physics sessions at the 2025 Global Physics Summit. The Collection is being guest edited by David Schwab (CUNY, New York) and Yuhai Tu (IBM Watson Research Center Yorktown Heights, NY). Every article published in this collection underwent a rigorous peer review process, adhering to the same high standards applied to all papers. The Physical Review E editorial team managed the peer review and made all editorial decisions.
Viewpoint
Viewing Neural Networks Through a Statistical-Physics Lens
Statistical physics is shedding light on how network architecture and data structure shape the effectiveness of neural-network learning.
See more in Physics
Article Text
References (33)
- G. E. Hinton, N. Srivastava, A. Krizhevsky, I. Sutskever, and R. R. Salakhutdinov, Improving neural networks by preventing co-adaptation of feature detectors, arXiv:1207.0580.
- N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, Dropout: A simple way to prevent neural networks from overfitting, J. Mach. Learn. Res. 15, 1929 (2014).
- I. Salehin and D.-K. Kang, A review on dropout regularization approaches for deep neural networks within the scholarly domain, Electronics 12, 3106 (2023).
- L. Wan, M. Zeiler, S. Zhang, Y. Le Cun, and R. Fergus, Regularization of neural networks using dropconnect, in Proceedings of the 30th International Conference on Machine Learning (PMLR, 2013), pp. 1058–1066.
- P. Morerio, J. Cavazza, R. Volpi, R. Vidal, and V. Murino, in Proceedings of the IEEE International Conference on Computer Vision (IEEE, Piscataway, 2017), pp. 3544–3552.
- Z. Liu, Z. Xu, J. Jin, Z. Shen, and T. Darrell, in Proceedings of the 40th International Conference on Machine Learning (PMLR, Cambridge, 2023), pp. 22233–22248.
- W. Mou, Y. Zhou, J. Gao, and L. Wang, in Proceedings of the 35th International Conference on Machine Learning, edited by J. Dy and A. Krause (PMLR, Cambridge, 2018), Vol. 80, pp. 3645–3653.
- R. Arora, P. Bartlett, P. Mianjy, and N. Srebro, in Proceedings of the 38th International Conference on Machine Learning (PMLR, Cambridge, 2021), pp. 351–361.
- K. Zhai and H. Wang, in Proceedings of the Sixth International Conference on Learning Representations, Vancouver, 2018 (ICLR, La Jolla, 2018).
- P. Baldi and P. J. Sadowski, in Advances in Neural Information Processing Systems 26, edited by C. Burges, L. Bottou, M. Welling, Z. Ghahramani, and K. Weinberger (Curran, Red Hook, 2013).
- S. Wager, S. Wang, and P. Liang, Dropout training as adaptive regularization, in Advances in Neural Information Processing Systems 26 (Curran Associates, Red Hook, 2013), pp. 351–359.
- P. Mianjy, R. Arora, and R. Vidal, On the implicit bias of dropout, in Proceedings of the 35th International Conference on Machine Learning (PMLR, 2018), Vol. 80, pp. 3540–3548.
- Z. Zhang and Z.-Q. J. Xu, Implicit regularization of dropout, IEEE Trans. Pattern Anal. Mach. Intell. 46, 4206 (2024).
- S. S. Schoenholz, J. Gilmer, S. Ganguli, and J. Sohl-Dickstein, in Proceedings of the Fifth International Conference on Learning Representations, Toulon, 2017 (ICLR, La Jolla, 2017).
- M. Biehl and H. Schwarze, Learning by on-line gradient descent, J. Phys. A: Math. Gen. 28, 643 (1995).
- D. Saad and S. A. Solla, Exact solution for on-line learning in multilayer neural networks, Phys. Rev. Lett. 74, 4337 (1995).
- D. Saad and S. A. Solla, On-line learning in soft committee machines, Phys. Rev. E 52, 4225 (1995).
- S. Goldt, M. Advani, A. M. Saxe, F. Krzakala, and L. Zdeborová, in Advances in Neural Information Processing Systems 32, Vancouver, 2019, edited by H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett (Curran, Red Hook, 2019).
- E. Gardner and B. Derrida, Three unfinished works on the optimal storage capacity of networks, J. Phys. A: Math. Gen. 22, 1983 (1989).
- H. S. Seung, H. Sompolinsky, and N. Tishby, Statistical mechanics of learning from examples, Phys. Rev. A 45, 6056 (1992).
- A. Engel, Statistical Mechanics of Learning (Cambridge University Press, Cambridge, 2001).
- For full details on the calculations see the Mathematica notebooks at https://github.com/francescomori/analytic_dropout.git (2025).
- S. Goldt, M. Mézard, F. Krzakala, and L. Zdeborová, Modeling the influence of data structure on learning in neural networks: The hidden manifold model, Phys. Rev. X 10, 041044 (2020).
- M. Biehl, P. Riegler, and C. Wöhler, Transient dynamics of on-line learning in two-layered neural networks, J. Phys. A: Math. Gen. 29, 4769 (1996).
- M. Straat and M. Biehl, On-line learning dynamics of ReLu neural networks using statistical physics techniques, arXiv:1903.07378.
- D. Jarvis, S. Lee, C. C. J. Dominé, A. M. Saxe, and S. S. Mannelli, in The 13th International Conference on Learning Representations, Singapore, 2025 (ICLR, La Jolla, 2025).
- F. Richert, R. Worschech, and B. Rosenow, Soft mode in the dynamics of over-realizable online learning for soft committee machines, Phys. Rev. E 105, L052302 (2022).
- E. Oostwal, M. Straat, and M. Biehl, Hidden unit specialization in layered neural networks: ReLu vs. sigmoidal activation, Physica A 564, 125517 (2021).
- B. Aubin, A. Maillard, J. Barbier, F. Krzakala, N. Macris, L. Zdeborová, in Advances in Neural Information Processing Systems 31, Montreal, 2018, edited by S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (Curran, Red Hook, 2018).
- E. Agoritsas, G. Biroli, P. Urbani, and F. Zamponi, Out-of-equilibrium dynamical mean-field equations for the perceptron model, J. Phys. A: Math. Theor. 51, 085002 (2018).
- F. Mignacco, F. Krzakala, P. Urbani, and L. Zdeborová, in Advances in Neural Information Processing Systems 33, 2020, edited by H. Larochelle, M. Ranzato, R. Hadsell, M.-F. Balcan, and H.-T. Lin (Curran, Red Hook, 2020), p. 9540.
- F. Mori, S. S. Mannelli, and F. Mignacco, Optimal protocols for continual learning via statistical physics and control theory, J. Stat. Mech. (2025) 084004.
- F. Mignacco and F. Mori, A statistical physics framework for optimal learning, arXiv:2507.07907.