- Letter
- Open Access
Asymptotic generalization errors in the online learning of random feature models
Phys. Rev. Research 6, L022049 – Published 3 June, 2024
DOI: https://doi.org/10.1103/PhysRevResearch.6.L022049
Abstract
Deep neural networks are widely used prediction algorithms whose performance often improves as the number of weights increases, leading to over-parametrization. We consider a two-layered neural network whose first layer is frozen while the last layer is trainable, known as the random feature model. We study over-parametrization in the context of a student-teacher framework by deriving a set of differential equations for the learning dynamics. For any finite ratio of hidden layer size and input dimension, the student cannot generalize perfectly, and we compute the non-zero asymptotic generalization error. Only when the student's hidden layer size is exponentially larger than the input dimension, an approach to perfect generalization is possible.
Physics Subject Headings (PhySH)
Article Text
Supplemental Material
References (52)
- Y. LeCun, Y. Bengio, and G. Hinton, Deep learning, Nature (London) 521, 436 (2015).
- I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning (MIT Press, Cambridge, MA, 2016).
- A. Krizhevsky, I. Sutskever, and G. E. Hinton, Imagenet classification with deep convolutional neural networks, in Advances in Neural Information Processing Systems, edited by F. Pereira, C. Burges, L. Bottou, and K. Weinberger (Curran Associates, Inc., Glasgow, Scotland, 2012), Vol. 25.
- D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, Y. Chen, T. Lillicrap, F. Hui, L. Sifre, G. van den Driessche, T. Graepel, and D. Hassabis, Mastering the game of go without human knowledge, Nature (London) 550, 354 (2017).
- G. Carleo, I. Cirac, K. Cranmer, L. Daudet, M. Schuld, N. Tishby, L. Vogt-Maranto, and L. Zdeborová, Machine learning and the physical sciences, Rev. Mod. Phys. 91, 045002 (2019).
- C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals, Understanding deep learning (still) requires rethinking generalization, Commun. ACM 64, 107 (2021).
- M. Theunissen, M. Davel, and E. Barnard, Benign interpolation of noise in deep learning, South African Comput. J. 32, 12 (2020).
- A. Canziani, A. Paszke, and E. Culurciello, An analysis of deep neural network models for practical applications, arXiv:1605.07678 (2017).
- B. Neyshabur, Z. Li, S. Bhojanapalli, Y. LeCun, and N. Srebro, The role of over-parametrization in generalization of neural networks, in International Conference on Learning Representations (Curran Associates, Inc., Glasgow, Scotland, 2023).
- P. Bartlett, The sample complexity of pattern classification with neural networks: The size of the weights is more important than the size of the network, IEEE Trans. Inf. Theory 44, 525 (1998).
- R. M. Neal, Priors for Infinite Networks (Springer, New York, 1996), p. 29.
- J. Lee, J. Sohl-Dickstein, J. Pennington, R. Novak, S. Schoenholz, and Y. Bahri, Deep neural networks as gaussian processes, in International Conference on Learning Representations (2018).
- R. Novak, L. Xiao, Y. Bahri, J. Lee, G. Yang, D. A. Abolafia, J. Pennington, and J. Sohl-Dickstein, Bayesian deep convolutional networks with many channels are gaussian processes, in International Conference on Learning Representations (2019).
- A. G. de G. Matthews, J. Hron, M. Rowland, R. E. Turner, and Z. Ghahramani, Gaussian process behaviour in wide deep neural networks, in International Conference on Learning Representations (2018).
- C. Williams, Computing with infinite networks, in Advances in Neural Information Processing Systems, edited by M. Mozer, M. Jordan, and T. Petsche (MIT Press, Cambridge, MA, 1996), Vol. 9.
- A. Jacot, F. Gabriel, and C. Hongler, Neural tangent kernel: Convergence and generalization in neural networks, in Advances in Neural Information Processing Systems, edited by S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (Curran Associates, Inc., Glasgow, Scotland, 2018), Vol. 31.
- S. Arora, S. S. Du, W. Hu, Z. Li, R. R. Salakhutdinov, and R. Wang, On exact computation with an infinitely wide neural net, in Advances in Neural Information Processing Systems, edited by H. Wallach, H. Larochelle, A. Beygelzimer, F. d' Alché-Buc, E. Fox, and R. Garnett (Curran Associates, Inc., Glasgow, Scotland, 2019), Vol. 32.
- Z. Chen, Y. Cao, Q. Gu, and T. Zhang, A generalized neural tangent kernel analysis for two-layer neural networks, in Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, edited by H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin (Curran Associates, Inc., Glasgow, Scotland, 2021).
- O. Cohen, O. Malka, and Z. Ringel, Learning curves for overparametrized deep neural networks: A field theory perspective, Phys. Rev. Res. 3, 023034 (2021).
- S. Arora, S. Du, W. Hu, Z. Li, and R. Wang, Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks, in Proceedings of the 36th International Conference on Machine Learning, edited by K. Chaudhuri and R. Salakhutdinov, Vol. 97 of Proceedings of Machine Learning Research (Curran Associates, Inc., Glasgow, Scotland, 2019), pp. 322–332.
- L. Chizat, E. Oyallon, and F. Bach, On lazy training in differentiable programming, in Advances in Neural Information Processing Systems, edited by, H. Wallach, H. Larochelle, A. Beygelzimer, F. d' Alché-Buc, E. Fox, and R. Garnett (Curran Associates, Glasgow, Scotland, 2019), Vol. 32.
- J. Lee, L. Xiao, S. Schoenholz, Y. Bahri, R. Novak, J. Sohl-Dickstein, and J. Pennington, Wide neural networks of any depth evolve as linear models under gradient descent, in Advances in Neural Information Processing Systems, edited by H. Wallach, H. Larochelle, A. Beygelzimer, F. d' Alché-Buc, E. Fox, and R. Garnett (Curran Associates, Inc., Glasgow, Scotland, 2019), Vol. 32.
- A. Rahimi and B. Recht, Random features for large-scale kernel machines, in Advances in Neural Information Processing Systems, edited by, J. Platt, D. Koller, Y. Singer, and S. Roweis (Curran Associates, Inc., Glasgow, Scotland, 2007), Vol. 20.
- E. Gardner and B. Derrida, Three unfinished works on the optimal storage capacity of networks, J. Phys. A 22, 1983 (1989).
- J. Hertz, A. Krough, and R. Palmer, Introduction To The Theory Of Neural Computation (West View Press, Boulder, Colorado, 1991), Vol. 44, p. 12.
- T. L. H. Watkin, A. Rau, and M. Biehl, The statistical mechanics of learning a rule, Rev. Mod. Phys. 65, 499 (1993).
- D. Saad, On-line learning in neural networks, J. Am. Stat. Assoc. 95, 452 (2000).
- A. Engel and C. Van den Broeck, Statistical Mechanics of Learning (Cambridge University Press, Cambridge, 2001).
- Y. Bahri, J. Kadmon, J. Pennington, S. S. Schoenholz, J. Sohl-Dickstein, and S. Ganguli, Statistical mechanics of deep learning, Annu. Rev. Condens. Matter Phys. 11, 501 (2020).
- E. Gardner, The space of interactions in neural network models, J. Phys. A 21, 257 (1988).
- H. S. Seung, H. Sompolinsky, and N. Tishby, Statistical mechanics of learning from examples, Phys. Rev. A 45, 6056 (1992).
- P. Riegler and M. Biehl, On-line backpropagation in two-layered neural networks, J. Phys. A 28, L507 (1995).
- S. Goldt, M. Advani, A. M. Saxe, F. Krzakala, and L. Zdeborová, Dynamics of stochastic gradient descent for two-layer neural networks in the teacher-student setup, in Advances in Neural Information Processing Systems, edited by H. Wallach, H. Larochelle, A. Beygelzimer, F. d' Alché-Buc, E. Fox, and R. Garnett (Curran Associates, Inc., Glasgow, Scotland, 2019), Vol. 32.
- M. Biehl and H. Schwarze, Learning by on-line gradient descent, J. Phys. A. 28, 643 (1995).
- D. Saad and S. A. Solla, On-line learning in soft committee machines, Phys. Rev. E 52, 4225 (1995).
- D. Saad and S. A. Solla, Exact solution for on-line learning in multilayer neural networks, Phys. Rev. Lett. 74, 4337 (1995).
- M. Biehl, P. Riegler, and C. Wöhler, Transient dynamics of on-line learning in two-layered neural networks, J. Phys. A 29, 4769 (1996).
- M. Biehl, E. Schlösser, and M. Ahr, Phase transitions in soft-committee machines, Europhys. Lett. 44, 261 (1998).
- E. Oostwal, M. Straat, and M. Biehl, Hidden unit specialization in layered neural networks: Relu vs. sigmoidal activation, Physica A 564, 125517 (2021).
- F. Richert, R. Worschech, and B. Rosenow, Soft mode in the dynamics of over-realizable online learning for soft committee machines, Phys. Rev. E 105, L052302 (2022).
- G. Naveh, O. Ben David, H. Sompolinsky, and Z. Ringel, Predicting the outputs of finite deep neural networks trained with noisy gradients, Phys. Rev. E 104, 064301 (2021).
- Q. Li and H. Sompolinsky, Statistical mechanics of deep linear neural networks: The backpropagating kernel renormalization, Phys. Rev. X 11, 031059 (2021).
- G. Yehudai and O. Shamir, On the Power and Limitations of Random Features for Understanding Neural Networks (Curran Associates Inc., Glasgow, Scotland, 2019).
- S. Mei, T. Misiakiewicz, and A. Montanari, Generalization error of random feature and kernel methods: Hypercontractivity and kernel matrix concentration, Appl. Comput. Harmon. Anal. 59, 3 (2022), Special Issue on Harmonic Analysis and Machine Learning.
- B. Ghorbani, S. Mei, T. Misiakiewicz, and A. Montanari, Linearized two-layers neural networks in high dimension, Ann. Statist. 49, 1029 (2021).
- B. Ghorbani, S. Mei, T. Misiakiewicz, and A. Montanari, Limitations of lazy training of two-layers neural network, in Advances in Neural Information Processing Systems, edited by H. Wallach, H. Larochelle, A. Beygelzimer, F. d' Alché-Buc, E. Fox, and R. Garnett (Curran Associates, Inc., Glasgow, Scotland, 2019), Vol. 32.
- See Supplemental Material at http://link.aps.org/supplemental/10.1103/PhysRevResearch.6.L022049 for the setup of differential equations, the derivation of the Langevin equation and the the asymptotic generalization error. The analytical results are futher supported by numerical investigations.
- G. Reents and R. Urbanczik, Self-averaging and on-line learning, Phys. Rev. Lett. 80, 5445 (1998).
- G. Rotskoff and E. Vanden-Eijnden, Trainability and accuracy of artificial neural networks: An interacting particle system approach, Commun. Pure Appl. Math. 75, 1889 (2022).
- N. V. Kampen, Stochastic Processes in Physics and Chemistry (North Holland, Amsterdam, 2007).
- V. A. Marčenko and L. A. Pastur, Distribution of eigenvalues for some sets of random matrices, Math. USSR-Sbornik 1, 457 (1967).
- R. B. Ash, Information Theory (Dover, New York, 1965).