- Open Access
Statistical Physics of Deep Learning: Optimal Learning of a Multilayer Perceptron near Interpolation
Phys. Rev. X 16, 031014 – Published 22 July, 2026
DOI: https://doi.org/10.1103/56sb-pdh6
Abstract
For four decades, statistical physics has been providing a framework to analyze neural networks. A long-standing question remained on its capacity to tackle deep-learning models capturing rich feature-learning effects, thus going beyond the narrow networks or kernel methods analyzed until now. We positively answer through the study of the supervised learning of a multilayer perceptron. Importantly, (i) its width scales as the input dimension, making it more prone to feature learning than ultrawide networks and more expressive than narrow ones or ones with fixed embedding layers, and (ii) we focus on the challenging interpolation regime where the number of trainable parameters and data are comparable, which forces the model to adapt to the task. We consider the matched teacher-student setting. Therefore, we provide the fundamental limits of learning random deep neural-network targets and identify the sufficient statistics describing what is learned by an optimally trained network as the data budget increases. A rich phenomenology emerges with various learning transitions. With enough data, optimal performance is attained through the model’s “specialization” toward the target, but it can be hard to reach for training algorithms, which get attracted by suboptimal solutions predicted by the theory. Specialization occurs inhomogeneously across layers, propagating from shallow toward deep ones but also across neurons in each layer. Furthermore, deeper targets are harder to learn. Despite its simplicity, the Bayes-optimal setting provides insights into how the depth, nonlinearity, and finite (proportional) width influence neural networks in the feature-learning regime that are potentially relevant in much more general settings.
Physics Subject Headings (PhySH)
Popular Summary
Deep neural networks dominate modern machine learning, yet the mechanisms behind feature extraction and generalization from data have resisted analytical description beyond toy models. Using statistical physics and random matrix theory, we obtain the Bayes-optimal generalization error for deep fully connected networks, and describe their behavior as the training set grows: The network learns by first fitting the data as a polynomial, and then specializes its units to specific features while deciphering the hidden rule. Moreover, specialization propagates from shallow toward deeper, more abstract layers. The result is a quantitatively detailed closed-form picture of feature learning in deep networks, and a controlled starting point for extending such analyses to architectures with stronger inductive biases.
Article Text
References (230)
- P. L. Bartlett, A. Montanari, and A. Rakhlin, Deep learning: A statistical viewpoint, Acta Numer. 30, 87 (2021).
- Y. LeCun, Y. Bengio, and G. Hinton, Deep learning, Nature (London) 521, 436 (2015).
- D. J. Amit, H. Gutfreund, and H. Sompolinsky, Spin-glass models of neural networks, Phys. Rev. A 32, 1007 (1985).
- E. Gardner, The space of interactions in neural network models, J. Phys. A 21, 257 (1988).
- E. Gardner and B. Derrida, Three unfinished works on the optimal storage capacity of networks, J. Phys. A 22, 1983 (1989).
- H. S. Seung, M. Opper, and H. Sompolinsky, Query by committee, in Proceedings of the Fifth Annual Workshop on Computational Learning Theory, COLT ’92 (Association for Computing Machinery, New York, USA, 1992), pp. 287–294.
- A. Engel, H. M. Köhler, F. Tschepke, H. Vollmayr, and A. Zippelius, Storage capacity and learning algorithms for two-layer neural networks, Phys. Rev. A 45, 7590 (1992).
- K. Kang, J.-H. Oh, C. Kwon, and Y. Park, Generalization in a two-layer neural network, Phys. Rev. E 48, 4805 (1993).
- D. O’Kane and O. Winther, Learning to classify in large committee machines, Phys. Rev. E 50, 3201 (1994).
- H. Schwarze and J. Hertz, Generalization in fully connected committee machines, Europhys. Lett. 21, 785 (1993).
- R. Urbanczik, Storage capacity of the fully-connected committee machine, J. Phys. A 30, L387 (1997).
- O. Winther, B. Lautrup, and J.-B. Zhang, Optimal learning in multilayer neural networks, Phys. Rev. E 55, 836 (1997).
- H. Schwarze and J. Hertz, Generalization in a large committee machine, Europhys. Lett. 20, 375 (1992).
- H. Schwarze, M. Opper, and W. Kinzel, Generalization in a two-layer neural network, Phys. Rev. A 46, R6185 (1992).
- G. Mato and N. Parga, Generalization properties of multilayered neural networks, J. Phys. A 25, 5047 (1992).
- R. Monasson and R. Zecchina, Weight space structure and internal representations: A direct approach to learning and generalization in multilayer neural networks, Phys. Rev. Lett. 75, 2432 (1995).
- B. Schottky, Phase transitions in the generalization behaviour of multilayer neural networks, J. Phys. A 28, 4515 (1995).
- A. Engel, Correlation of internal representations in feed-forward neural networks, J. Phys. A 29, L323 (1996).
- D. Malzahn, A. Engel, and I. Kanter, Storage capacity of correlated perceptrons, Phys. Rev. E 55, 7369 (1997).
- D. Malzahn and A. Engel, Correlations between hidden units in multilayer neural networks and replica symmetry breaking, Phys. Rev. E 60, 2097 (1999).
- H. Sompolinsky, N. Tishby, and H. S. Seung, Learning from examples in large neural networks, Phys. Rev. Lett. 65, 1683 (1990).
- G. Györgyi, First-order transition to perfect generalization in a neural network with binary synapses, Phys. Rev. A 41, 7097 (1990).
- R. Meir and J. F. Fontanari, Learning from examples in weight-constrained neural networks, J. Phys. A 25, 1149 (1992).
- D. M. L. Barbato and J. F. Fontanari, The effects of lesions on the generalization ability of a perceptron, J. Phys. A 26, 1847 (1993).
- A. Engel and L. Reimers, Reliability of replica symmetry for the generalization problem of a toy multilayer neural network, Europhys. Lett. 28, 531 (1994).
- G. J. Bex, R. Serneels, and C. Van den Broeck, Storage capacity and generalization error for the reversed-wedge Ising perceptron, Phys. Rev. E 51, 6309 (1995).
- E. Barkai, D. Hansel, and H. Sompolinsky, Broken symmetries in multilayered perceptrons, Phys. Rev. A 45, 4146 (1992).
- H. Schwarze, Learning a rule in a multilayer neural network, J. Phys. A 26, 5781 (1993).
- A. Engel and C. Van den Broeck, Statistical Mechanics of Learning (Cambridge University Press, Cambridge, England, 2001).
- H. Cui, High-dimensional learning of narrow neural networks, J. Stat. Mech. (2025) 023402.
- J. Bruna and D. Hsu, Survey on Algorithms for Multi-Index Models, Stat. Sci. 40, 378 (2025).
- G. B. Arous, R. Gheissari, and A. Jagannath, Online stochastic gradient descent on non-convex losses from high-dimensional inference, J. Mach. Learn. Res. 22 (2021).
- A. Damian, L. Pillaud-Vivien, J. Lee, and J. Bruna, Computational-statistical gaps in Gaussian single-index models (extended abstract), in Proceedings of Thirty Seventh Conference on Learning Theory, Proceedings of Machine Learning Research Vol. 247, edited by S. Agrawal and A. Roth (PMLR, 2024), pp. 1262–1262.
- E. Abbe, E. Boix-Adsera, M. Brennan, G. Bresler, and D. Nagaraj, The staircase property: How hierarchical structure can guide deep learning, in Proceedings of the 35th International Conference on Neural Information Processing Systems, NIPS ’21 (Curran Associates Inc., Red Hook, USA, 2021).
- E. Abbe, E. B. Adserà, and T. Misiakiewicz, SGD learning on neural networks: leap complexity and saddle-to-saddle dynamics, in Proceedings of Thirty Sixth Conference on Learning Theory, Proceedings of Machine Learning Research, Vol. 195, edited by G. Neu and L. Rosasco (PMLR, 2023), pp. 2552–2623.
- E. Troiani, Y. Dandi, L. Defilippis, L. Zdeborova, B. Loureiro, and F. Krzakala, Fundamental computational limits of weak learnability in high-dimensional multi-index models, in Proceedings of The 28th International Conference on Artificial Intelligence and Statistics, Proceedings of Machine Learning Research Vol. 258, edited by Y. Li, S. Mandt, S. Agrawal, and E. Khan (PMLR, 2025), pp. 2467–2475.
- R. M. Neal, Priors for infinite networks, in Bayesian Learning for Neural Networks (Springer, New York, 1996), pp. 29–53.
- C. Williams, Computing with infinite networks, in Advances in Neural Information Processing Systems, edited by M. Mozer, M. Jordan, and T. Petsche (MIT Press, Cambridge, MA, 1996), Vol. 9.
- J. Lee, J. Sohl-dickstein, J. Pennington, R. Novak, S. Schoenholz, and Y. Bahri, Deep neural networks as Gaussian processes, in International Conference on Learning Representations (2018).
- A. G. D. G. Matthews, J. Hron, M. Rowland, R. E. Turner, and Z. Ghahramani, Gaussian process behaviour in wide deep neural networks, in International Conference on Learning Representations (2018).
- B. Hanin, Random neural networks in the infinite width limit as Gaussian processes, Ann. Appl. Probab. 33, 4798 (2023).
- H. Yoon and J.-H. Oh, Learning of higher-order perceptrons with tunable complexities, J. Phys. A 31, 7771 (1998).
- R. Dietrich, M. Opper, and H. Sompolinsky, Statistical mechanics of support vector networks, Phys. Rev. Lett. 82, 2975 (1999).
- F. Gerace, B. Loureiro, F. Krzakala, M. Mézard, and L. Zdeborová, Generalisation error in learning with random features and the hidden manifold model, J. Stat. Mech. (2021) 124013.
- B. Bordelon, A. Canatar, and C. Pehlevan, Spectrum dependent learning curves in kernel regression and wide neural networks, in Proceedings of the 37th International Conference on Machine Learning, Proceedings of Machine Learning Research Vol. 119, edited by H. D. III and A. Singh (PMLR, 2020), pp. 1024–1034.
- A. Canatar, B. Bordelon, and C. Pehlevan, Spectral bias and task-model alignment explain generalization in kernel regression and infinitely wide neural networks, Nat. Commun. 12, 2914 (2021).
- L. Xiao, H. Hu, T. Misiakiewicz, Y. M. Lu, and J. Pennington, Precise learning curves and higher-order scaling limits for dot-product kernel regression, J. Stat. Mech. (2023) 114005.
- B. Ghorbani, S. Mei, T. Misiakiewicz, and A. Montanari, Linearized two-layers neural networks in high dimension, Ann. Stat. 49, 1029 (2021).
- A. Rahimi and B. Recht, Random features for large-scale kernel machines, in Advances in Neural Information Processing Systems, edited by J. Platt, D. Koller, Y. Singer, and S. Roweis (Curran Associates, Inc., 2007), Vol. 20.
- A. Jacot, F. Gabriel, and C. Hongler, Neural tangent kernel: Convergence and generalization in neural networks, in Advances in Neural Information Processing Systems, edited by S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (Curran Associates, Inc., 2018), Vol. 31.
- L. Chizat, E. Oyallon, and F. Bach, On lazy training in differentiable programming, in Advances in Neural Information Processing Systems, edited by H. Wallach, H. Larochelle, A. Beygelzimer, F. d’Alché-Buc, E. Fox, and R. Garnett (Curran Associates, Inc., 2019), Vol. 32.
- B. Ghorbani, S. Mei, T. Misiakiewicz, and A. Montanari, When do neural networks outperform kernel methods?, in Advances in Neural Information Processing Systems, edited by H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin (Curran Associates, Inc., 2020), Vol. 33, pp. 14820–14830.
- M. Refinetti, S. Goldt, F. Krzakala, and L. Zdeborova, Classifying high-dimensional Gaussian mixtures: Where kernel methods fail and neural networks succeed, in Proceedings of the 38th International Conference on Machine Learning, Proceedings of Machine Learning Research Vol. 139, edited by M. Meila and T. Zhang (PMLR, 2021), pp. 8936–8947.
- E. Dyer and G. Gur-Ari, Asymptotics of wide networks from Feynman diagrams, in International Conference on Learning Representations (2020).
- S. Yaida, Non-Gaussian processes and neural networks at finite widths, in Proceedings of The First Mathematical and Scientific Machine Learning Conference, Proceedings of Machine Learning Research Vol. 107, edited by J. Lu and R. Ward (PMLR, 2020), pp. 165–192.
- G. Naveh, O. Ben David, H. Sompolinsky, and Z. Ringel, Predicting the outputs of finite deep neural networks trained with noisy gradients, Phys. Rev. E 104, 064301 (2021).
- J. Zavatone-Veth, A. Canatar, B. Ruben, and C. Pehlevan, Asymptotics of representation learning in finite Bayesian neural networks, in Advances in Neural Information Processing Systems, edited by M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan (Curran Associates, Inc., 2021), Vol. 34, pp. 24765–24777.
- K. T. Grosvenor and R. Jefferson, The edge of chaos: Quantum field theory and deep neural networks, SciPost Phys. 12, 081 (2022).
- K. Fischer, J. Lindner, D. Dahmen, Z. Ringel, M. Krämer, and M. Helias, Critical feature learning in deep neural networks, in Proceedings of the 41st International Conference on Machine Learning, Proceedings of Machine Learning Research Vol. 235, edited by R. Salakhutdinov, Z. Kolter, K. Heller, A. Weller, N. Oliver, J. Scarlett, and F. Berkenkamp (PMLR, 2024), pp. 13660–13690.
- I. Banta, T. Cai, N. Craig, and Z. Zhang, Structures of neural network effective theories, Phys. Rev. D 109, 105007 (2024).
- M. Guillen, P. Misof, and J. E. Gerken, Finite-width neural tangent kernels from Feynman diagrams, arXiv:2508.11522.
- Y. Bahri, B. Hanin, A. Brossollet, V. Erba, C. Keup, R. Pacelli, and J. B. Simon, Les Houches lectures on deep learning at large and infinite width*, J. Stat. Mech. (2024) 104012.
- Z. Ringel, N. Rubin, E. Mor, M. Helias, and I. Seroussi, Applications of statistical field theory in deep learning, arXiv:2502.18553.
- S. Mei, A. Montanari, and P.-M. Nguyen, A mean field view of the landscape of two-layer neural networks, Proc. Natl. Acad. Sci. U.S.A. 115, E7665 (2018).
- S. Mei, T. Misiakiewicz, and A. Montanari, Mean-field theory of two-layers neural networks: Dimension-free bounds and kernel limit, in Proceedings of the Thirty-Second Conference on Learning Theory, Proceedings of Machine Learning Research Vol. 99, edited by A. Beygelzimer and D. Hsu (PMLR, 2019), pp. 2388–2464.
- G. Yang and E. J. Hu, Tensor programs IV: Feature learning in infinite-width neural networks, in Proceedings of the 38th International Conference on Machine Learning, Proceedings of Machine Learning Research Vol. 139, edited by M. Meila and T. Zhang (PMLR, 2021), pp. 11727–11737.
- G. Rotskoff and E. Vanden-Eijnden, Trainability and accuracy of artificial neural networks: An interacting particle system approach, Commun. Pure Appl. Math. 75, 1889 (2022).
- J. Sirignano and K. Spiliopoulos, Mean field analysis of neural networks: A central limit theorem, Stoch. Proc. Appl. 130, 1820 (2020).
- B. Bordelon and C. Pehlevan, Self-consistent dynamical field theory of kernel evolution in wide neural networks, in Advances in Neural Information Processing Systems, edited by S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Curran Associates, Inc., 2022), Vol. 35, pp. 32240–32256.
- P.-M. Nguyen and H. T. Pham, A rigorous framework for the mean field limit of multilayer neural networks, Math. Stat. Learn. 6, 201 (2023).
- I. Seroussi, G. Naveh, and Z. Ringel, Separation of scales and a thermodynamic description of feature learning in some CNNs, Nat. Commun. 14, 908 (2023).
- F. Bassetti, M. Gherardi, A. Ingrosso, M. Pastore, and P. Rotondo, Feature learning in finite-width Bayesian deep linear networks with multiple outputs and convolutional layers, J. Mach. Learn. Res. 26, 1 (2025).
- N. Rubin, Z. Ringel, I. Seroussi, and M. Helias, A unified approach to feature learning in Bayesian neural networks, in High-dimensional Learning Dynamics 2024: The Emergence of Structure and Reasoning (2024).
- A. van Meegen and H. Sompolinsky, Coding schemes in neural networks learning classification tasks, Nat. Commun. 16, 3354 (2025).
- C. Lauditi, B. Bordelon, and C. Pehlevan, Adaptive kernel predictors from feature-learning infinite limits of neural networks, arXiv:2502.07998.
- A. X. Yang, M. Robeyns, E. Milsom, B. Anson, N. Schoots, and L. Aitchison, A theory of representation learning gives a deep generalisation of kernel methods, in Proceedings of the 40th International Conference on Machine Learning, Proceedings of Machine Learning Research Vol. 202, edited by A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett (PMLR, 2023), pp. 39380–39415.
- N. Rubin, I. Seroussi, and Z. Ringel, Grokking as a first order phase transition in two layer networks, in The Twelfth International Conference on Learning Representations (2024).
- A. M. Saxe, J. McClelland, and S. Ganguli, Exact solutions to the nonlinear dynamics of learning in deep linear neural networks, in Proceedings of the International Conference on Learning Representations 2014 (2014).
- Q. Li and H. Sompolinsky, Statistical mechanics of deep linear neural networks: The backpropagating kernel renormalization, Phys. Rev. X 11, 031059 (2021).
- L. Aitchison, Why bigger is not always better: On finite and infinite neural networks, in International Conference on Machine Learning (PMLR, 2020), pp. 156–164.
- B. Hanin and A. Zlokapa, Bayesian interpolation with deep linear networks, Proc. Natl. Acad. Sci. U.S.A. 120, e2301345120 (2023).
- J. A. Zavatone-Veth, W. L. Tong, and C. Pehlevan, Contrasting random and learned features in deep Bayesian linear regression, Phys. Rev. E 105, 064118 (2022).
- B. Neyshabur, R. Tomioka, and N. Srebro, Norm-based capacity control in neural networks, in Proceedings of The 28th Conference on Learning Theory, Proceedings of Machine Learning Research Vol. 40, edited by P. Grünwald, E. Hazan, and S. Kale (PMLR, Paris, France, 2015), pp. 1376–1401.
- S. Pesme and N. Flammarion, Saddle-to-saddle dynamics in diagonal linear networks, in Advances in Neural Information Processing Systems, edited by A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Curran Associates, Inc., 2023), Vol. 36, pp. 7475–7505.
- D. Soudry, E. Hoffer, M. S. Nacson, S. Gunasekar, and N. Srebro, The implicit bias of gradient descent on separable data, J. Mach. Learn. Res. 19, 1 (2018).
- S. Pesme, L. Pillaud-Vivien, and N. Flammarion, Implicit bias of SGD for diagonal linear networks: A provable benefit of stochasticity, in Advances in Neural Information Processing Systems, edited by A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan (2021).
- R. Berthier, Incremental learning in diagonal linear networks, J. Mach. Learn. Res. 24, (2023), http://jmlr.org/papers/v24/22-1395.html.
- H. Labarrière, C. Molinari, L. Rosasco, S. Villa, and C. Vega, Optimization insights into deep diagonal linear networks, arXiv:2412.16765.
- S. Du and J. Lee, On the power of over-parametrization in neural networks with quadratic activation, in Proceedings of the 35th International Conference on Machine Learning, Proceedings of Machine Learning Research Vol. 80, edited by J. Dy and A. Krause (PMLR, 2018), pp. 1329–1338.
- M. Soltanolkotabi, A. Javanmard, and J. D. Lee, Theoretical insights into the optimization landscape of over-parameterized shallow neural networks, IEEE Trans. Inf. Theory 65, 742 (2019).
- L. Venturi, A. S. Bandeira, and J. Bruna, Spurious valleys in one-hidden-layer neural network optimization landscapes, J. Mach. Learn. Res. 20, 1 (2019).
- S. Sarao Mannelli, E. Vanden-Eijnden, and L. Zdeborová, Optimization and generalization of shallow neural networks with quadratic activation functions, in Advances in Neural Information Processing Systems, edited by H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin (Curran Associates, Inc., 2020), Vol. 33, pp. 13445–13455.
- D. Gamarnik, E. C. K𝚤z𝚤ldağ, and I. Zadik, Stationary points of a shallow neural network with quadratic activations and the global optimality of the gradient descent algorithm, Math. Oper. Res. 50, 209 (2024).
- S. Martin, F. Bach, and G. Biroli, On the impact of overparameterization on the training of a shallow neural network in high dimensions, in Proceedings of The 27th International Conference on Artificial Intelligence and Statistics, Proceedings of Machine Learning Research Vol. 238, edited by S. Dasgupta, S. Mandt, and Y. Li (PMLR, 2024), pp. 3655–3663.
- Y. Arjevani, J. Bruna, J. Kileel, E. Polak, and M. Trager, Geometry and optimization of shallow polynomial networks, arXiv:2501.06074.
- A. Maillard, E. Troiani, S. Martin, L. Zdeborová, and F. Krzakala, Bayes-optimal learning of an extensive-width neural network from quadratically many samples, in Advances in Neural Information Processing Systems, edited by A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang (Curran Associates, Inc., 2024), Vol. 37, pp. 82085–82132.
- Y. Xu, A. Maillard, L. Zdeborová, and F. Krzakala, Fundamental limits of matrix sensing: Exact asymptotics, universality, and applications (2025).
- V. Erba, E. Troiani, L. Zdeborová, and F. Krzakala, The nuclear route: Sharp asymptotics of ERM in overparameterized quadratic networks, arXiv:2505.17958.
- G. Ben Arous, M. A. Erdogdu, N. M. Vural, and D. Wu, Learning quadratic neural networks in high dimensions: SGD dynamics and scaling laws, arXiv:2508.03688.
- J. Barbier, F. Krzakala, N. Macris, L. Miolane, and L. Zdeborová, Optimal errors and phase transitions in high-dimensional generalized linear models, Proc. Natl. Acad. Sci. U.S.A. 116, 5451 (2019).
- J. Barbier and N. Macris, Statistical limits of dictionary learning: Random matrix theory and the spectral replica method, Phys. Rev. E 106, 024136 (2022).
- A. Maillard, F. Krzakala, M. Mézard, and L. Zdeborová, Perturbative construction of mean-field equations in extensive-rank matrix factorization and denoising, J. Stat. Mech. (2022) 083301.
- F. Pourkamali, J. Barbier, and N. Macris, Matrix inference in growing rank regimes, IEEE Trans. Inf. Theory 70, 8133 (2024).
- G. Semerjian, Matrix denoising: Bayes-optimal estimators via low-degree polynomials, J. Stat. Phys. 191, 139 (2024).
- R. Pacelli, S. Ariosto, M. Pastore, F. Ginelli, M. Gherardi, and P. Rotondo, A statistical mechanics framework for Bayesian deep neural networks beyond the infinite-width limit, Nat. Mach. Intell. 5, 1497 (2023).
- P. Baglioni, R. Pacelli, R. Aiudi, F. Di Renzo, A. Vezzani, R. Burioni, and P. Rotondo, Predictive power of a Bayesian effective action for fully connected one hidden layer neural networks in the proportional limit, Phys. Rev. Lett. 133, 027301 (2024).
- A. Ingrosso, R. Pacelli, P. Rotondo, and F. Gerace, Statistical mechanics of transfer learning in fully connected networks in the proportional limit, Phys. Rev. Lett. 134, 177301 (2025).
- B. Hanin and A. Zlokapa, Gibbs measures from deep shaped multilayer perceptrons, Phys. Rev. Lett. 136, 067301 (2026).
- H. Cui, F. Krzakala, and L. Zdeborova, Bayes-optimal learning of deep random networks of extensive-width, in Proceedings of the 40th International Conference on Machine Learning, Proceedings of Machine Learning Research Vol. 202, edited by A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett (PMLR, 2023), pp. 6468–6521.
- F. Camilli, D. Tieplova, and J. Barbier, Fundamental limits of overparametrized shallow neural networks for supervised learning, Boll. Unione Mat. Ital. (2025), 10.1007/s40574-025-00506-2.
- F. Camilli, D. Tieplova, E. Bergamin, and J. Barbier, Information-theoretic reduction of deep neural networks to linear models in the overparametrized proportional regime, in Proceedings of Thirty Eighth Conference on Learning Theory, Proceedings of Machine Learning Research Vol. 291, edited by N. Haghtalab and A. Moitra (PMLR, 2025), pp. 757–798.
- G. Naveh and Z. Ringel, A self consistent theory of Gaussian processes captures feature learning effects in finite CNNs, in Advances in Neural Information Processing Systems, edited by M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan (Curran Associates, Inc., 2021), Vol. 34, pp. 21352–21364.
- R. Aiudi, R. Pacelli, P. Baglioni, A. Vezzani, R. Burioni, and P. Rotondo, Local kernel renormalization as a mechanism for feature learning in overparametrized convolutional neural networks, Nat. Commun. 16, 568 (2025).
- H. Yoshino, From complex to simple: Hierarchical free-energy landscape renormalized in deep neural networks, SciPost Phys. Core 2, 005 (2020).
- H. Yoshino, Spatially heterogeneous learning by a deep student machine, Phys. Rev. Res. 5, 033068 (2023).
- G. Huang, L. S. Chan, H. Yoshino, G. Zhang, and Y. Jin, Liquid and solid layers in a thermal deep learning machine, arXiv:2506.06789.
- J. Yao, Y. Yacoby, B. Coker, W. Pan, and F. Doshi-Velez, An empirical analysis of the advantages of finite- v.s. infinite-width Bayesian neural networks (2022).
- J. Lee, S. S. Schoenholz, J. Pennington, B. Adlam, L. Xiao, R. Novak, and J. Sohl-Dickstein, Finite versus infinite neural networks: An empirical study, in Proceedings of the 34th International Conference on Neural Information Processing Systems, NIPS ’20 (Curran Associates Inc., Red Hook, NY, USA, 2020).
- L. Zdeborová, Understanding deep learning is also a job for physicists, Nat. Phys. 16, 602 (2020).
- Y. Bahri, J. Kadmon, J. Pennington, S. S. Schoenholz, J. Sohl-Dickstein, and S. Ganguli, Statistical mechanics of deep learning, Annu. Rev. Condens. Matter Phys. 11, 501 (2020).
- J. Hoffmann et al., Training compute-optimal large language models, in Advances in Neural Information Processing Systems, edited by S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Curran Associates, Inc., 2022), Vol. 35, pp. 30016–30030.
- J. Ni, Q. Liu, C. Du, L. Dou, H. Yan, Z. Wang, T. Pang, and M. Q. Shieh, Training optimal large diffusion language models, arXiv:2510.03280.
- M. Lan, P. Torr, A. Meek, A. Khakzar, D. Krueger, and F. Barez, Quantifying feature space universality across large language models via sparse autoencoders, arXiv:2410.06981.
- Z. Li, C. Fan, and T. Zhou, Grokking in LLM pretraining? Monitor memorization-to-generalization without test, arXiv:2506.21551.
- M. Mondelli and A. Montanari, On the connection between learning two-layer neural networks and tensor decomposition, in Proceedings of the Twenty-Second International Conference on Artificial Intelligence and Statistics, Proceedings of Machine Learning Research Vol. 89, edited by K. Chaudhuri and M. Sugiyama (PMLR, 2019), pp. 1051–1060.
- M. Mezard, G. Parisi, and M. Virasoro, Spin Glass Theory and Beyond (World Scientific, Singapore, 1986).
- C. Itzykson and J. Zuber, The planar approximation. II, J. Math. Phys. (N.Y.) 21, 411 (1980).
- A. Matytsin, On the large- limit of the Itzykson-Zuber integral, Nucl. Phys. B411, 805 (1994).
- A. Guionnet and O. Zeitouni, Large deviations asymptotics for spherical integrals, J. Funct. Anal. 188, 461 (2002).
- A. Guionnet, First order asymptotics of matrix integrals; a rigorous approach towards the understanding of matrix models, Commun. Math. Phys. 244, 527 (2004).
- J.-B. Zuber, The large- limit of matrix integrals over the orthogonal group, J. Phys. A 41, 382001 (2008).
- V. A. Kazakov, Solvable matrix models, arXiv:hep-th/0003064.
- E. Brézin, S. Hikami et al., Random Matrix Theory with an External Source (Springer, New York, 2016).
- D. Anninos and B. Mühlmann, Notes on matrix models (matrix musings), J. Stat. Mech. (2020) 083109.
- J. Bun, J. P. Bouchaud, S. N. Majumdar, and M. Potters, Instanton approach to large Harish-Chandra-Itzykson-Zuber integrals, Phys. Rev. Lett. 113, 070201 (2014).
- M. Potters and J.-P. Bouchaud, A First Course in Random Matrix Theory: For Physicists, Engineers and Data Scientists (Cambridge University Press, Cambridge, England, 2020).
- J. Husson and J. Ko, Spherical integrals of sublinear rank, Probab. Theory Relat. Fields 193, 1 (2025).
- G. Parisi and M. Potters, Mean-field equations for spin models with orthogonal interaction matrices, J. Phys. A 28, 5267 (1995).
- M. Opper and O. Winther, Adaptive and self-averaging Thouless-Anderson-Palmer mean-field theory for probabilistic modeling, Phys. Rev. E 64, 056131 (2001).
- M. Opper, B. Çakmak, and O. Winther, A theory of solving TAP equations for Ising models with general invariant random matrices, J. Phys. A 49, 114002 (2016).
- Z. Fan, Y. Li, and S. Sen, TAP equations for orthogonally invariant spin glasses at high temperature, arXiv:2202.09325.
- J. Barbier and M. Sáenz, Marginals of a spherical spin glass model with correlated disorder, Electron. Commun. Probab. 27, 1 (2022).
- Z. Fan and Y. Wu, The replica-symmetric free energy for Ising spin glasses with orthogonally invariant couplings, Probab. Theory Relat. Fields 190, 1 (2024).
- Y. Kabashima, Inference from correlated patterns: A unified theory for perceptron learning and linear vector channels, J. Phys. Conf. Ser. 95, 012001 (2008).
- M. Gabrié, A. Manoel, C. Luneau, J. Barbier, N. Macris, F. Krzakala, and L. Zdeborová, Entropy and mutual information in models of deep neural networks, in Advances in Neural Information Processing Systems, edited by S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (Curran Associates, Inc., 2018), Vol. 31.
- K. Takeda, S. Uda, and Y. Kabashima, Analysis of CDMA systems that are characterized by eigenvalue spectrum, Europhys. Lett. 76, 1193 (2006).
- A. Tulino, G. Caire, S. Shamai, and S. Verdú, Support recovery with sparsely sampled free random matrices, in 2011 IEEE International Symposium on Information Theory Proceedings (2011), pp. 2328–2332.
- T. Hou, Y. Liu, T. Fu, and J. Barbier, Sparse superposition codes under VAMP decoding with generic rotational invariant coding matrices, in 2022 IEEE International Symposium on Information Theory (ISIT) (2022), pp. 1372–1377.
- J. Barbier, F. Camilli, Y. Xu, and M. Mondelli, Information limits and Thouless-Anderson-Palmer equations for spiked matrix models with structured noise, Phys. Rev. Res. 7, 013081 (2025).
- S. Rangan, P. Schniter, and A. K. Fletcher, Vector approximate message passing, IEEE Trans. Inf. Theory 65, 6664 (2019).
- J. Ma and L. Ping, Orthogonal AMP, IEEE Access 5, 2020 (2017).
- A. Maillard, L. Foini, A. L. Castellanos, F. Krzakala, M. Mézard, and L. Zdeborová, High-temperature expansions and message passing algorithms, J. Stat. Mech. (2019) 113301.
- L. Liu, S. Huang, and B. M. Kurkoski, Memory AMP, IEEE Trans. Inf. Theory 68, 8015 (2022).
- K. Takeuchi, On the convergence of orthogonal/vector AMP: Long-memory message-passing strategy, in 2022 IEEE International Symposium on Information Theory (ISIT) (2022), pp. 1366–1371.
- T. Takahashi and Y. Kabashima, Macroscopic analysis of vector approximate message passing in a model-mismatched setting, IEEE Trans. Inf. Theory 68, 5579 (2022).
- Z. Fan, Approximate Message Passing algorithms for rotationally invariant matrices, Ann. Stat. 50, 197 (2022).
- J. Barbier, N. Macris, A. Maillard, and F. Krzakala, The mutual information in random linear estimation beyond i.i.d. matrices, in 2018 IEEE International Symposium on Information Theory (ISIT) (2018), pp. 1390–1394.
- C. Gerbelot, A. Abbara, and F. Krzakala, Asymptotic errors for high-dimensional convex penalized linear regression beyond Gaussian matrices, in Proceedings of Thirty Third Conference on Learning Theory, Proceedings of Machine Learning Research Vol. 125, edited by J. Abernethy and S. Agarwal (PMLR, 2020), pp. 1682–1713.
- C. Gerbelot, A. Abbara, and F. Krzakala, Asymptotic errors for teacher-student convex generalized linear models (or: How to prove Kabashima’s replica formula), IEEE Trans. Inf. Theory 69, 1824 (2023).
- R. Dudeja, Y. M. Lu, and S. Sen, Universality of approximate message passing with semirandom matrices, Ann. Probab. 51, 1616 (2023).
- J. Barbier, F. Camilli, M. Mondelli, and M. Sáenz, Fundamental limits in structured principal component analysis and how to reach them, Proc. Natl. Acad. Sci. U.S.A. 120, e2302028120 (2023).
- R. Dudeja, S. Liu, and J. Ma, Optimality of approximate message passing algorithms for spiked matrix models with rotationally invariant noise, arXiv:2405.18081.
- O. Ledoit and S. Péché, Eigenvectors of some large sample covariance matrix ensembles, Probab. Theory Relat. Fields 151, 233 (2011).
- J. Bun, R. Allez, J.-P. Bouchaud, and M. Potters, Rotational invariant estimator for general noisy matrices, IEEE Trans. Inf. Theory 62, 7475 (2016).
- F. Pourkamali and N. Macris, Rectangular rotational invariant estimator for general additive noise matrices, in 2023 IEEE International Symposium on Information Theory (ISIT) (2023), pp. 2081–2086.
- E. Troiani, V. Erba, F. Krzakala, A. Maillard, and L. Zdeborova, Optimal denoising of rotationally invariant rectangular matrices, in Proceedings of Mathematical and Scientific Machine Learning, Proceedings of Machine Learning Research, Vol. 190, edited by B. Dong, Q. Li, L. Wang, and Z.-Q. J. Xu (PMLR, 2022), pp. 97–112.
- H. C. Schmidt, Statistical physics of sparse and dense models in optimization and inference, Ph.D. thesis, IPHT—Institut de Physique Théorique, 2018.
- A. Sakata and Y. Kabashima, Statistical mechanics of dictionary learning, Europhys. Lett. 103, 28008 (2013).
- Y. Kabashima, F. Krzakala, M. Mézard, A. Sakata, and L. Zdeborová, Phase transitions and sample complexity in Bayes-optimal matrix factorization, IEEE Trans. Inf. Theory 62, 4228 (2016).
- V. Erba, E. Troiani, L. Biggio, A. Maillard, and L. Zdeborová, Bilinear sequence regression: A model for learning from long sequences of high-dimensional tokens, Phys. Rev. X 15, 021092 (2025).
- J. Barbier, F. Camilli, J. Ko, and K. Okajima, Phase diagram of extensive-rank symmetric matrix denoising beyond rotational invariance, Phys. Rev. X 15, 021085 (2025).
- Y. Ren, E. Nichani, D. Wu, and J. D. Lee, Emergence and scaling laws in SGD learning of shallow neural networks, arXiv:2504.19983.
- A. Bodin and N. Macris, Gradient flow on extensive-rank positive semi-definite matrix denoising, in 2023 IEEE Information Theory Workshop (ITW) (IEEE, 2023), pp. 365–370.
- M.-T. Nguyen and R. Skerk, Statistical physics of deep learning (2026), https://github.com/Minh-Toan/statphys-deep-NN.
- S. Mei and A. Montanari, The generalization error of random features regression: Precise asymptotics and the double descent curve, Commun. Pure Appl. Math. 75, 667 (2022).
- S. Goldt, M. Mézard, F. Krzakala, and L. Zdeborová, Modeling the influence of data structure on learning in neural networks: The hidden manifold model, Phys. Rev. X 10, 041044 (2020).
- T. Hastie, A. Montanari, S. Rosset, and R. J. Tibshirani, Surprises in high-dimensional ridgeless least squares interpolation, Ann. Stat. 50, 949 (2022).
- S. Goldt, B. Loureiro, G. Reeves, F. Krzakala, M. Mezard, and L. Zdeborová, The Gaussian equivalence of generative models for learning with shallow neural networks, in Proceedings of the 2nd Mathematical and Scientific Machine Learning Conference, Proceedings of Machine Learning Research Vol. 145, edited by J. Bruna, J. Hesthaven, and L. Zdeborová (PMLR, 2022), pp. 426–471.
- H. Hu and Y. M. Lu, Universality laws for high-dimensional learning with random features, IEEE Trans. Inf. Theory 69, 1932 (2023).
- A. Montanari and B. N. Saeed, Universality of empirical risk minimization, in Proceedings of Thirty Fifth Conference on Learning Theory, Proceedings of Machine Learning Research Vol. 178, edited by P.-L. Loh and M. Raginsky (PMLR, 2022), pp. 4310–4312,
- G. G. Wen, H. Hu, Y. M. Lu, Z. Fan, and T. Misiakiewicz, When does Gaussian equivalence fail and how to fix it: Non-universal behavior of random features with quadratic scaling, arXiv:2512.03325.
- H. Nishimori, Statistical Physics of Spin Glasses and Information Processing: An Introduction (Oxford University Press, New York, 2001).
- L. Zdeborová and F. Krzakala, Statistical physics of inference: Thresholds and algorithms, Adv. Phys. 65, 453 (2016).
- D. Guo, S. Shamai, and S. Verdú, Mutual information and minimum mean-square error in Gaussian channels, IEEE Trans. Inf. Theory 51, 1261 (2005).
- A. Guionnet and J. Huang, Asymptotics of rectangular spherical integrals, J. Funct. Anal. 285, 110144 (2023).
- J. Barbier and D. Panchenko, Strong replica symmetry in high-dimensional optimal Bayesian inference, Commun. Math. Phys. 393, 1199 (2022).
- J. T. Parker, P. Schniter, and V. Cevher, Bilinear generalized approximate message passing—Part I: Derivation, IEEE Trans. Signal Process. 62, 5839 (2014).
- F. Krzakala, M. Mézard, and L. Zdeborová, Phase diagram and approximate message passing for blind calibration and dictionary learning, in 2013 IEEE International Symposium on Information Theory (2013), pp. 659–663.
- B. Aubin, A. Maillard, J. Barbier, F. Krzakala, N. Macris, and L. Zdeborová, The committee machine: Computational to statistical gaps in learning a two-layers neural network, in Advances in Neural Information Processing Systems, edited by S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (Curran Associates, Inc., 2018), Vol. 31.
- C. Baldassi, E. M. Malatesta, and R. Zecchina, Properties of the geometry of solutions and capacity of multilayer neural networks with rectified linear unit activations, Phys. Rev. Lett. 123, 170602 (2019).
- J. Barbier, F. Gerace, A. Ingrosso, C. Lauditi, E. M. Malatesta, G. Nwemadji, and R. P. Ortiz, Generalization performance of narrow one-hidden layer networks in the teacher-student setting, arXiv:2507.00629.
- J. Barbier, Overlap matrix concentration in optimal Bayesian inference, Inf. Inference 10, 597 (2020).
- A. Maillard, E. Troiani, S. Martin, F. Krzakala, and L. Zdeborová, Github repository ExtensiveWidthQuadraticSamples, https://github.com/SPOC-group/ExtensiveWidthQuadraticSamples (2024).
- T. Tao and V. Vu, Random matrices: Universality of local eigenvalue statistics up to the edge, Commun. Math. Phys. 298, 549 (2010).
- D. P. Kingma and J. Ba, Adam: A method for stochastic optimization, arXiv:1412.6980.
- M. Hennick and S. D. Baerdemacker, Almost Bayesian: The fractal dynamics of stochastic gradient descent, arXiv:2503.22478.
- C. Mingard, G. Valle-Pérez, J. Skalse, and A. A. Louis, Is SGD a Bayesian sampler? Well, almost, J. Mach. Learn. Res. 22, 1 (2021).
- S. L. Smith, D. Duckworth, S. Rezchikov, Q. V. Le, and J. Sohl-Dickstein, Stochastic natural gradient descent draws posterior samples in function space, arXiv:1806.09597.
- S. Mandt, M. D. Hoffman, and D. M. Blei, Stochastic gradient descent as approximate Bayesian inference, J. Mach. Learn. Res. 18, 1 (2017).
- M. Raghu, J. Gilmer, J. Yosinski, and J. Sohl-Dickstein, SVCCA: Singular vector canonical correlation analysis for deep learning dynamics and interpretability, in Advances in Neural Information Processing Systems, edited by I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Curran Associates, Inc., 2017), Vol. 30.
- F. Cagnetta, L. Petrini, U. M. Tomasini, A. Favero, and M. Wyart, How deep neural networks learn compositional data: The random hierarchy model, Phys. Rev. X 14, 031001 (2024).
- F. Aguirre-López, S. Franz, and M. Pastore, Random features and polynomial rules, SciPost Phys. 18, 039 (2025).
- H. Hu, Y. M. Lu, and T. Misiakiewicz, Asymptotics of random feature regression beyond the linear scaling regime, arXiv:2403.08160.
- J. Barbier and N. Macris, The adaptive interpolation method: A simple scheme to prove replica formulas in Bayesian inference, Probab. Theory Relat. Fields 174, 1133 (2019).
The work [206] concurrent with ours proposes an annealed approximation of the free entropy (rather than the more accurate quenched computation proposed in the present paper) to study shallow MLPs of extensive width with ReLU activation.
- A. Afanah and B. Rosenow, Unified description of learning dynamics in the soft committee machine from finite to ultra-wide regimes, arXiv:2512.16556.
- R. Monasson, Properties of neural networks storing spatially correlated patterns, J. Phys. A 25, 3701 (1992).
- B. Loureiro, C. Gerbelot, H. Cui, S. Goldt, F. Krzakala, M. Mezard, and L. Zdeborová, Learning curves of generic features maps for realistic datasets with a teacher-student model, in Advances in Neural Information Processing Systems, edited by M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan (Curran Associates, Inc., 2021), Vol. 34, pp. 18137–18151.
- Del Giudice, P., Franz, S., and M. A. Virasoro, Perceptron beyond the limit of capacity, J. Phys. (Les Ulis, Fr.) 50, 121 (1989).
- B. Loureiro, G. Sicuro, C. Gerbelot, A. Pacco, F. Krzakala, and L. Zdeborová, Learning Gaussian mixtures with generalized linear models: Precise asymptotics in high-dimensions, in Advances in Neural Information Processing Systems, edited by M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan (Curran Associates, Inc., 2021), Vol. 34, pp. 10144–10157.
- B. Lopez, M. Schroder, and M. Opper, Storage of correlated patterns in a perceptron, J. Phys. A 28, L447 (1995).
- S. Y. Chung, D. D. Lee, and H. Sompolinsky, Classification and geometry of general perceptual manifolds, Phys. Rev. X 8, 031003 (2018).
- P. Rotondo, M. Pastore, and M. Gherardi, Beyond the storage capacity: Data-driven satisfiability transition, Phys. Rev. Lett. 125, 120601 (2020).
- M. Pastore, P. Rotondo, V. Erba, and M. Gherardi, Statistical learning theory of structured data, Phys. Rev. E 102, 032119 (2020).
- A. Sclocchi, A. Favero, and M. Wyart, A phase transition in diffusion models reveals the hierarchical nature of data, Proc. Natl. Acad. Sci. U.S.A. 122, e2408799121 (2025).
- D. Saad and S. A. Solla, On-line learning in soft committee machines, Phys. Rev. E 52, 4225 (1995).
- D. Saad and S. Solla, Dynamics of on-line gradient descent learning for multilayer neural networks, in Advances in Neural Information Processing Systems, edited by D. Touretzky, M. Mozer, and M. Hasselmo (MIT Press, Cambridge, MA, 1995), Vol. 8.
- S. Goldt, M. S. Advani, A. M. Saxe, F. Krzakala, and L. Zdeborová, Dynamics of stochastic gradient descent for two-layer neural networks in the teacher-student setup, J. Stat. Mech. (2020) 124010.
- L. F. Cugliandolo, Recent applications of dynamical mean-field methods, Annu. Rev. Condens. Matter Phys. 15, 177 (2024).
- A. Montanari and P. Urbani, Dynamical decoupling of generalization and overfitting in large two-layer networks, arXiv:2502.21269.
- B. Bordelon, A. Atanasov, and C. Pehlevan, A dynamical model of neural scaling laws, in Proceedings of the 41st International Conference on Machine Learning, Proceedings of Machine Learning Research Vol. 235, edited by R. Salakhutdinov, Z. Kolter, K. Heller, A. Weller, N. Oliver, J. Scarlett, and F. Berkenkamp (PMLR, 2024), pp. 4345–4382.
- E. Paquette, C. Paquette, L. Xiao, and J. Pennington, phases of compute-optimal neural scaling laws, in Advances in Neural Information Processing Systems, edited by A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang (Curran Associates, Inc., 2024), Vol. 37, pp. 16459–16537.
- L. Lin, J. Wu, S. M. Kakade, P. L. Bartlett, and J. D. Lee, Scaling laws in linear regression: Compute, parameters, and data, in Advances in Neural Information Processing Systems, edited by A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang (Curran Associates, Inc., 2024), Vol. 37, pp. 60556–60606.
- K. Oko, Y. Song, T. Suzuki, and D. Wu, Learning sum of diverse features: Computational hardness and efficient gradient-based training for ridge combinations, in Proceedings of Thirty Seventh Conference on Learning Theory, Proceedings of Machine Learning Research, Vol. 247, edited by S. Agrawal and A. Roth (PMLR, 2024), pp. 4009–4081.
- A. Power, Y. Burda, H. Edwards, I. Babuschkin, and V. Misra, Grokking: Generalization beyond overfitting on small algorithmic datasets, arXiv:2201.02177.
- N. Rubin, O. Davidovich, and Z. Ringel, Mitigating the curse of detail: Scaling arguments for feature learning and sample complexity, arXiv:2512.04165.
- T. M. Cover, Elements of Information Theory (John Wiley & Sons, New York, 1999).
- M. Abadi et al., tensorflow: Large-scale machine learning on heterogeneous systems (2015), software available from tensorflow.org.
- M. D. Hoffman and A. Gelman, The No-U-Turn Sampler: Adaptively setting path lengths in Hamiltonian Monte Carlo, J. Mach. Learn. Res. 15, 1593 (2014).
- E. Bingham et al., Pyro: Deep universal probabilistic programming, J. Mach. Learn. Res. 20, 28:1 (2019).
