Reuse & Permissions

It is not necessary to obtain permission to reuse this article or its components as it is available under the terms of the Creative Commons Attribution 4.0 International license. This license permits unrestricted use, distribution, and reproduction in any medium, provided attribution to the author(s) and the published article's title, journal citation, and DOI are maintained. Please note that some figures may have been included with permission from other third parties. It is your responsibility to obtain the proper permission from the rights holder directly for these figures.

Export citation

Export citation

Choose format for download:

Download Citation
  • Letter
  • Open Access

Leveraging chaotic transients in the training of artificial neural networks

Pedro Jiménez-González, Miguel C. Soriano, and Lucas Lacasa*

  • *Contact author: lucas@ifisc.uib-csic.es

Phys. Rev. Research 8, L022032 – Published 18 May, 2026

DOI: https://doi.org/10.1103/t5p9-kv5w

Abstract

Traditional algorithms to optimize artificial neural networks when confronted with a supervised learning task are usually exploitation-type relaxational dynamics such as gradient descent (GD). Here, we explore the dynamics of the neural network trajectory along training for unconventionally large learning rates. We show that for a region of values of the learning rate, the GD optimization shifts away from purely exploitationlike algorithm into a regime of exploration-exploitation balance, as the neural network is still capable of learning but the trajectory shows sensitive dependence on initial conditions—as characterized by positive network maximum Lyapunov exponent. Interestingly, the characteristic training time required to reach an acceptable accuracy in the test set reaches a minimum precisely in such learning rate region, further suggesting that one can accelerate the training of artificial neural networks by locating at the onset of chaos. Our results—initially illustrated for the MNIST classification task—qualitatively hold for a range of supervised learning tasks, learning architectures (including both shallow and deep multilayer perceptrons and convolutional neural networks) and other hyperparameters (different activation functions and weight regularization), and showcase the emergent, constructive role of transient chaotic dynamics in the training of artificial neural networks.

View figure in article

Physics Subject Headings (PhySH)

Article Text

References (52)

  1. I. J. Goodfellow, Y. Bengio, and A. Courville, Deep Learning (MIT Press, Cambridge, MA, USA, 2016), http://www.deeplearningbook.org.
  2. C. C. Aggarwal, Neural Networks and Deep Learning (Springer, Cham, 2018), Vol. 10.
  3. L. Lacasa, J. P. Rodriguez, and V. M. Eguiluz, Correlations of network trajectories, Phys. Rev. Res. 4, L042008 (2022).
  4. L. A. N. Amaral, Artificial intelligence needs a scientific method-driven reset, Nat. Phys. 20, 523 (2024).
  5. G. Bianconi, A. Arenas, J. Biamonte, L. D Carr, B. Kahng, J. Kertesz, J. Kurths, L. Lü, C. Masoller, A. E. Motter, et al., Complex systems in the spotlight: Next steps after the 2021 Nobel Prize in Physics, J. Phys.: Complexity 4, 010201 (2023).
  6. L. Arola-Fernández and L. Lacasa, Effective theory of collective deep learning, Phys. Rev. Res. 6, L042040 (2024).
  7. K. Danovski, M. C. Soriano, and L. Lacasa, Dynamical stability and chaos in artificial neural network trajectories along training, Front. Complex Syst. 2, 1367957 (2024).
  8. V. Latora, V. Nicosia, and G. Russo, Complex Networks: Principles, Methods and Applications (Cambridge University Press, Cambridge, 2017).
  9. P. Holme and J. Saramäki, Temporal networks, Phys. Rep. 519, 97 (2012).
  10. N. Masuda and R. Lambiotte, A Guide to Temporal Networks (World Scientific, Singapore, 2016).
  11. O. E. Williams, L. Lacasa, A. P. Millán, and V. Latora, The shape of memory in temporal networks, Nat. Commun. 13, 499 (2022).
  12. A. Badie-Modiri, C. Boldrini, L. Valerio, J. Kertész, and M. Karsai, Initialisation and network effects in decentralised federated learning, Appl. Network Sci. 10, 53 (2025).
  13. E. L. Malfa, G. L. Malfa, G. Nicosia, and V. Latora, Deep neural networks via complex network theory: A perspective, in Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI) (2024), pp. 4361–4369.
  14. Z. Zheng, H. Liang, V. Snasel, V. Latora, P. Pardalos, G. Nicosia, and V. Ojha, On learnable parameters of optimal and suboptimal deep learning models, in Proceedings of the International Conference on Neural Information Processing (Springer, 2024), pp. 126–140.
  15. H. G. Schuster, W. Just, Deterministic Chaos: An Introduction (John Wiley & Sons, Weinheim, 2006).
  16. H. Kantz and T. Schreiber, Nonlinear Time Series Analysis (Cambridge University Press, Cambridge, 2003).
  17. D. Bertsekas, A. Nedic, and A. Ozdaglar, Convex Analysis and Optimization (Athena Scientific, Belmont, MA, 2003), Vol. 1.
  18. L. Kong and M. Tao, Stochasticity of deterministic gradient descent: Large learning rate for multiscale objective function, Adv. Neural Inf. Process. Syst. 33, 2625 (2020).
  19. L. Herrmann, M. Granz, and T. Landgraf, Chaotic dynamics are intrinsic to neural network training with SGD, Adv. Neural Inf. Process. Syst. 35, 5219 (2022).
  20. Y.-C. Lai and T. Tél, Transient Chaos: Complex Dynamics on Finite Time Scales (Springer Science & Business Media, New York, 2011), Vol. 173.
  21. M. Črepinšek, S.-H. Liu, and M. Mernik, Exploration and exploitation in evolutionary algorithms: A survey, ACM Comput. Surv. (CSUR) 45, 35 (2013).
  22. O. Berger-Tal, J. Nathan, E. Meron, and D. Saltz, The exploration-exploitation dilemma: A multidisciplinary framework, PLoS One 9, e95693 (2014).
  23. N. E. Humphries, H. Weimerskirch, N. Queiroz, E. J. Southall, and D. W. Sims, Foraging success of biological Lévy flights recorded in situ, Proc. Natl. Acad. Sci. USA 109, 7169 (2012).
  24. G. Ramos-Fernández, J. L. Mateos, O. Miramontes, G. Cocho, H. Larralde, and B. Ayala-Orozco, Lévy walk patterns in the foraging movements of spider monkeys (Ateles geoffroyi), Behav. Ecol. Sociobiol. 55, 223 (2004).
  25. S. Eliassen, C. Jørgensen, M. Mangel, and J. Giske, Exploration or exploitation: Life expectancy changes the value of learning in foraging strategies, Oikos 116, 513 (2007).
  26. A. Reynolds, E. Ceccon, C. Baldauf, T. K. Medeiros, and O. Miramontes, Lévy foraging patterns of rural humans, PLoS One 13, e0199099 (2018).
  27. J. M Kembro, M. Lihoreau, J. Garriga, E. P Raposo, and F. Bartumeus, Bumblebees learn foraging routes through exploitation–exploration cycles, J. R. Soc. Interface 16, 20190103 (2019).
  28. C. T. Monk, M. Barbier, P. Romanczuk, J. R. Watson, J. Alós, S. Nakayama, D. I. Rubenstein, S. A. Levin, and R. Arlinghaus, How ecology shapes exploitation: A framework to predict the behavioural response of human and animal foragers along exploration–exploitation trade-offs, Ecol. Lett. 21, 779 (2018).
  29. L. R. Paiva, S. G. Alves, L. Lacasa, O. DeSouza, and O. Miramontes, Visibility graphs of animal foraging trajectories, J. Phys.: Complexity 3, 04LT03 (2022).
  30. R. Klages, G. Radons, and I. M. Sokolov, Anomalous Transport: Foundations and Applications (Wiley-VCH, Weinheim, 2008).
  31. T. T. Hills, P. M. Todd, D. Lazer, A. D. Redish, and I. D. Couzin, Exploration versus exploitation in space, mind, and society, Trends Cognit. Sci. 19, 46 (2015).
  32. M. A. Addicott, J. M. Pearson, M. M. Sweitzer, D. L. Barack, and M. L. Platt, A primer on foraging and the explore/exploit trade-off for psychiatry research, Neuropsychopharmacology 42, 1931, (2017).
  33. S. Ishii, W. Yoshida, and J. Yoshimoto, Control of exploitation–exploration meta-parameter in reinforcement learning, Neural Networks 15, 665 (2002).
  34. N. Bertschinger, T. Natschläger, and R. Legenstein, At the edge of chaos: Real-time computations and self-organized criticality in recurrent neural networks, Adv. Neural Inf. Process. Syst. 17, 145 (2004).
  35. J. Kadmon and H. Sompolinsky, Transition to chaos in random neuronal networks, Phys. Rev. X 5, 041030 (2015).
  36. D. Sussillo and L. F. Abbott, Generating coherent patterns of activity from chaotic neural networks, Neuron 63, 544 (2009).
  37. U. Pereira-Obilinovic, J. Aljadeff, and N. Brunel, Forgetting leads to chaos in attractor networks, Phys. Rev. X 13, 011009 (2023).
  38. D. Pazó, Discontinuous transition to chaos in a canonical random neural network, Phys. Rev. E 110, 014201 (2024).
  39. J. Luo, J. Chen, and H.-K. Zhang, The butterfly effect in neural networks: Unveiling hyperbolic chaos through parameter sensitivity, Neural Networks 189, 107572 (2025).
  40. L. Storm, H. Linander, J. Bec, K. Gustavsson, and B. Mehlig, Finite-time Lyapunov exponents of deep neural networks, Phys. Rev. Lett. 132, 057301 (2024).
  41. J. Cohen, S. Kaur, Y. Li, J. Z. Kolter, and A. Talwalkar, Gradient descent on neural networks typically occurs at the edge of stability, in International Conference on Learning Representations (2021).
  42. C. G. Langton, Computation at the edge of chaos: Phase transitions and emergent computation, Physica D 42, 12 (1990).
  43. Y. LeCun, The MNIST database of handwritten digits (1998), http://yann.lecun.com/exdb/mnist/.
  44. A. Caligiuri, V. M. Eguíluz, L. Di Gaetano, T. Galla, and L. Lacasa, Lyapunov exponents for temporal networks, Phys. Rev. E 107, 044305 (2023).
  45. A. Caligiuri, T. Galla, and L. Lacasa, Characterizing the dynamics of unlabeled temporal networks, Chaos 35, 053122 (2025).
  46. P. F. M. J. Verschure, Chaos-based learning, Complex Syst. 5, 359 (1991).
  47. Y. Dandi, F. Krzakala, B. Loureiro, L. Pesce, and L. Stephan, How two-layer neural networks learn, one (giant) step at a time, J. Mach. Learn. Res. 25, 1 (2024).
  48. L. Lacasa, F. J. Marín-Rodríguez, N. Masuda, and L. Arola-Fernández, Scalar embedding of temporal network trajectories, Chaos, Solitons Fractals 199, 116599 (2025).
  49. L. Lacasa, Fluid dynamics meet network science: Two cases of temporal network eigendecomposition, arXiv:2509.03135.
  50. https://github.com/pedrojg8/chaotic-transients-in-ANNs.
  51. R. A. Fisher, The use of multiple measurements in taxonomic problems, Ann. Eugen. 7, 179 (1936).
  52. A. Krizhevsky and G. Hinton, Learning multiple layers of features from tiny images, Master's thesis, Department of Computer Science, University of Toronto, 2009.

Outline

Information

Sign In to Your Journals Account

Filter

Filter

Article Lookup

Enter a citation