Reuse & Permissions

It is not necessary to obtain permission to reuse this article or its components as it is available under the terms of the Creative Commons Attribution 4.0 International license. This license permits unrestricted use, distribution, and reproduction in any medium, provided attribution to the author(s) and the published article's title, journal citation, and DOI are maintained. Please note that some figures may have been included with permission from other third parties. It is your responsibility to obtain the proper permission from the rights holder directly for these figures.

Export citation

Export citation

Choose format for download:

Download Citation
  • Letter
  • Open Access

Emergence of hierarchical modes from deep learning

Chan Li1 and Haiping Huang1,2,*

  • 1PMI Laboratory, School of Physics, Sun Yat-sen University, Guangzhou 510275, People's Republic of China
  • 2Guangdong Provincial Key Laboratory of Magnetoelectric Physics and Devices, Sun Yat-sen University, Guangzhou 510275, People's Republic of China

  • *huanghp7@mail.sysu.edu.cn

Phys. Rev. Research 5, L022011 – Published 11 April, 2023

DOI: https://doi.org/10.1103/PhysRevResearch.5.L022011

Abstract

Large-scale deep neural networks consume expensive training costs, but the training results in less-interpretable weight matrices constructing the networks. Here, we propose a mode decomposition learning that can interpret the weight matrices as a hierarchy of latent modes. These modes are akin to patterns in physics studies of memory networks, but the least number of modes increases only logarithmically with the network width and even becomes a constant when the width grows further. The mode decomposition learning not only saves a significant large amount of training costs but also explains the network performance with the leading modes, displaying a striking piecewise power-law behavior. The modes specify a progressively compact latent space across the network hierarchy, making a more disentangled subspace compared to standard training. Our mode decomposition learning is also studied in an analytic online learning setting, which reveals multiple stages of learning dynamics with a continuous specialization of hidden nodes. Therefore the proposed mode decomposition learning points to a cheap and interpretable route towards the magical deep learning.

View figure in article

Physics Subject Headings (PhySH)

Article Text

Supplemental Material

References (18)

  1. I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning (MIT Press, Cambridge, MA, 2016).
  2. G. Carleo, I. Cirac, K. Cranmer, L. Daudet, M. Schuld, N. Tishby, L. Vogt-Maranto, and L. Zdeborová, Machine learning and the physical sciences, Rev. Mod. Phys. 91, 045002 (2019).
  3. H. Huang, Statistical Mechanics of Neural Networks (Springer, Singapore, 2022).
  4. D. A. Roberts, S. Yaida, and B. Hanin, The Principles of Deep Learning Theory: An Effective Theory Approach to Understanding Neural Networks (Cambridge University Press, Cambridge, 2022).
  5. M. Jaderberg, A. Vedaldi, and A. Zisserman, Speeding up convolutional neural networks with low rank expansions, arXiv:1405.3866.
  6. H. Yang, M. Tang, W. Wen, F. Yan, D. Hu, A. Li, H. Li, and Y. Chen, Learning low-rank deep neural networks via singular vector orthogonality regularization and singular value sparsification, in Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) (IEEE, New York, 2020), pp. 2899–2908.
  7. L. Giambagli, L. Buffoni, T. Carletti, W. Nocentini, and D. Fanelli, Machine learning in spectral domain, Nat. Commun. 12, 1330 (2021).
  8. L. Chicchi, L. Giambagli, L. Buffoni, T. Carletti, M. Ciavarella, and D. Fanelli, Training of sparse and dense deep neural networks: Fewer parameters, same performance, Phys. Rev. E 104, 054312 (2021).
  9. Z. Jiang, J. Zhou, T. Hou, K. Y. M. Wong, and H. Huang, Associative memory model with arbitrary Hebbian length, Phys. Rev. E 104, 064306 (2021).
  10. J. Zhou, Z. Jiang, T. Hou, Z. Chen, K. Y. M. Wong, and H. Huang, Eigenvalue spectrum of neural networks with arbitrary Hebbian length, Phys. Rev. E 104, 064307 (2021).
  11. Y. LeCun, The MNIST database of handwritten digits, retrieved from http://yann.lecun.com/exdb/mnist.
  12. K. Fischer, A. René, C. Keup, M. Layer, D. Dahmen, and M. Helias, Decomposing neural networks as mappings of correlation functions, Phys. Rev. Res. 4, 043143 (2022).
  13. See Supplemental Material at http://link.aps.org/supplemental/10.1103/PhysRevResearch.5.L022011 for technical and experimental details. Codes are available at https://github.com/Chan-Li/MDL-model.
  14. M. Biehl and H. Schwarze, Learning by on-line gradient descent, J. Phys. A: Math. Gen. 28, 643 (1995).
  15. D. Saad and S. A. Solla, Exact Solution for On-Line Learning in Multilayer Neural Networks, Phys. Rev. Lett. 74, 4337 (1995).
  16. S. Goldt, M. S. Advani, A. M. Saxe, F. Krzakala, and L. Zdeborová, Generalisation dynamics of online learning in over-parameterised neural networks, arXiv:1901.09085.
  17. A. Bjoerck and G. H. Golub, Numerical methods for computing angles between linear subspaces, Math. Comput. 27, 579 (1973).
  18. H. Schwarze, Learning a rule in a multilayer neural network, J. Phys. A: Math. Gen. 26, 5781 (1993).

Outline

Information

Sign In to Your Journals Account

Filter

Filter

Article Lookup

Enter a citation