Reuse & Permissions

It is not necessary to obtain permission to reuse this article or its components as it is available under the terms of the Creative Commons Attribution 4.0 International license. This license permits unrestricted use, distribution, and reproduction in any medium, provided attribution to the author(s) and the published article's title, journal citation, and DOI are maintained. Please note that some figures may have been included with permission from other third parties. It is your responsibility to obtain the proper permission from the rights holder directly for these figures.

Export citation

Export citation

Choose format for download:

Download Citation
  • Featured in Physics
  • Open Access

Analytic theory of dropout regularization

Francesco Mori*

Francesca Mignacco

  • *Contact author: francesco.mori@physics.ox.ac.uk

Phys. Rev. E 112, 045301 – Published 1 October, 2025

DOI: https://doi.org/10.1103/jmdx-x3gr

Abstract

Dropout is a regularization technique widely used in training artificial neural networks to mitigate overfitting. It consists of dynamically deactivating subsets of the network during training to promote more robust representations. Despite its widespread adoption, dropout rates are often selected heuristically, and theoretical explanations of its success remain sparse. Here we analytically study dropout in two-layer neural networks trained with online stochastic gradient descent. In the high-dimensional limit, we derive a set of ordinary differential equations that fully characterize the evolution of the network during training and capture the effects of dropout. We obtain a number of exact results describing the generalization error and the optimal dropout probability at short, intermediate, and long training times. Our analysis shows that dropout reduces detrimental correlations between hidden nodes and mitigates the impact of label noise and that the optimal dropout probability increases with the level of noise in the data. Our results are validated by extensive numerical simulations.

View figure in article

Physics Subject Headings (PhySH)

Collections

This article appears in the following collection:

Statistical Physics Meets Machine Learning - Machine Learning Meets Statistical Physics

The Editors of Physical Review E are pleased to present the Collection on Statistical Physics Meets Machine Learning - Machine Learning Meets Statistical Physics, highlighting research at the intersection of machine learning and statistical physics, on the occasion of the two Statistical Physics Meets Machine Learning and the two Machine Learning Meets Statistical Physics sessions at the 2025 Global Physics Summit. The Collection is being guest edited by David Schwab (CUNY, New York) and Yuhai Tu (IBM Watson Research Center Yorktown Heights, NY). Every article published in this collection underwent a rigorous peer review process, adhering to the same high standards applied to all papers. The Physical Review E editorial team managed the peer review and made all editorial decisions.

Viewpoint

Viewing Neural Networks Through a Statistical-Physics Lens

Published 23 February, 2026

Statistical physics is shedding light on how network architecture and data structure shape the effectiveness of neural-network learning.

See more in Physics

Article Text

References (33)

  1. G. E. Hinton, N. Srivastava, A. Krizhevsky, I. Sutskever, and R. R. Salakhutdinov, Improving neural networks by preventing co-adaptation of feature detectors, arXiv:1207.0580.
  2. N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, Dropout: A simple way to prevent neural networks from overfitting, J. Mach. Learn. Res. 15, 1929 (2014).
  3. I. Salehin and D.-K. Kang, A review on dropout regularization approaches for deep neural networks within the scholarly domain, Electronics 12, 3106 (2023).
  4. L. Wan, M. Zeiler, S. Zhang, Y. Le Cun, and R. Fergus, Regularization of neural networks using dropconnect, in Proceedings of the 30th International Conference on Machine Learning (PMLR, 2013), pp. 1058–1066.
  5. P. Morerio, J. Cavazza, R. Volpi, R. Vidal, and V. Murino, in Proceedings of the IEEE International Conference on Computer Vision (IEEE, Piscataway, 2017), pp. 3544–3552.
  6. Z. Liu, Z. Xu, J. Jin, Z. Shen, and T. Darrell, in Proceedings of the 40th International Conference on Machine Learning (PMLR, Cambridge, 2023), pp. 22233–22248.
  7. W. Mou, Y. Zhou, J. Gao, and L. Wang, in Proceedings of the 35th International Conference on Machine Learning, edited by J. Dy and A. Krause (PMLR, Cambridge, 2018), Vol. 80, pp. 3645–3653.
  8. R. Arora, P. Bartlett, P. Mianjy, and N. Srebro, in Proceedings of the 38th International Conference on Machine Learning (PMLR, Cambridge, 2021), pp. 351–361.
  9. K. Zhai and H. Wang, in Proceedings of the Sixth International Conference on Learning Representations, Vancouver, 2018 (ICLR, La Jolla, 2018).
  10. P. Baldi and P. J. Sadowski, in Advances in Neural Information Processing Systems 26, edited by C. Burges, L. Bottou, M. Welling, Z. Ghahramani, and K. Weinberger (Curran, Red Hook, 2013).
  11. S. Wager, S. Wang, and P. Liang, Dropout training as adaptive regularization, in Advances in Neural Information Processing Systems 26 (Curran Associates, Red Hook, 2013), pp. 351–359.
  12. P. Mianjy, R. Arora, and R. Vidal, On the implicit bias of dropout, in Proceedings of the 35th International Conference on Machine Learning (PMLR, 2018), Vol. 80, pp. 3540–3548.
  13. Z. Zhang and Z.-Q. J. Xu, Implicit regularization of dropout, IEEE Trans. Pattern Anal. Mach. Intell. 46, 4206 (2024).
  14. S. S. Schoenholz, J. Gilmer, S. Ganguli, and J. Sohl-Dickstein, in Proceedings of the Fifth International Conference on Learning Representations, Toulon, 2017 (ICLR, La Jolla, 2017).
  15. M. Biehl and H. Schwarze, Learning by on-line gradient descent, J. Phys. A: Math. Gen. 28, 643 (1995).
  16. D. Saad and S. A. Solla, Exact solution for on-line learning in multilayer neural networks, Phys. Rev. Lett. 74, 4337 (1995).
  17. D. Saad and S. A. Solla, On-line learning in soft committee machines, Phys. Rev. E 52, 4225 (1995).
  18. S. Goldt, M. Advani, A. M. Saxe, F. Krzakala, and L. Zdeborová, in Advances in Neural Information Processing Systems 32, Vancouver, 2019, edited by H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett (Curran, Red Hook, 2019).
  19. E. Gardner and B. Derrida, Three unfinished works on the optimal storage capacity of networks, J. Phys. A: Math. Gen. 22, 1983 (1989).
  20. H. S. Seung, H. Sompolinsky, and N. Tishby, Statistical mechanics of learning from examples, Phys. Rev. A 45, 6056 (1992).
  21. A. Engel, Statistical Mechanics of Learning (Cambridge University Press, Cambridge, 2001).
  22. For full details on the calculations see the Mathematica notebooks at https://github.com/francescomori/analytic_dropout.git (2025).
  23. S. Goldt, M. Mézard, F. Krzakala, and L. Zdeborová, Modeling the influence of data structure on learning in neural networks: The hidden manifold model, Phys. Rev. X 10, 041044 (2020).
  24. M. Biehl, P. Riegler, and C. Wöhler, Transient dynamics of on-line learning in two-layered neural networks, J. Phys. A: Math. Gen. 29, 4769 (1996).
  25. M. Straat and M. Biehl, On-line learning dynamics of ReLu neural networks using statistical physics techniques, arXiv:1903.07378.
  26. D. Jarvis, S. Lee, C. C. J. Dominé, A. M. Saxe, and S. S. Mannelli, in The 13th International Conference on Learning Representations, Singapore, 2025 (ICLR, La Jolla, 2025).
  27. F. Richert, R. Worschech, and B. Rosenow, Soft mode in the dynamics of over-realizable online learning for soft committee machines, Phys. Rev. E 105, L052302 (2022).
  28. E. Oostwal, M. Straat, and M. Biehl, Hidden unit specialization in layered neural networks: ReLu vs. sigmoidal activation, Physica A 564, 125517 (2021).
  29. B. Aubin, A. Maillard, J. Barbier, F. Krzakala, N. Macris, L. Zdeborová, in Advances in Neural Information Processing Systems 31, Montreal, 2018, edited by S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (Curran, Red Hook, 2018).
  30. E. Agoritsas, G. Biroli, P. Urbani, and F. Zamponi, Out-of-equilibrium dynamical mean-field equations for the perceptron model, J. Phys. A: Math. Theor. 51, 085002 (2018).
  31. F. Mignacco, F. Krzakala, P. Urbani, and L. Zdeborová, in Advances in Neural Information Processing Systems 33, 2020, edited by H. Larochelle, M. Ranzato, R. Hadsell, M.-F. Balcan, and H.-T. Lin (Curran, Red Hook, 2020), p. 9540.
  32. F. Mori, S. S. Mannelli, and F. Mignacco, Optimal protocols for continual learning via statistical physics and control theory, J. Stat. Mech. (2025) 084004.
  33. F. Mignacco and F. Mori, A statistical physics framework for optimal learning, arXiv:2507.07907.

Outline

Information

Sign In to Your Journals Account

Filter

Filter

Article Lookup

Enter a citation