Reuse & Permissions

It is not necessary to obtain permission to reuse this article or its components as it is available under the terms of the Creative Commons Attribution 4.0 International license. This license permits unrestricted use, distribution, and reproduction in any medium, provided attribution to the author(s) and the published article's title, journal citation, and DOI are maintained. Please note that some figures may have been included with permission from other third parties. It is your responsibility to obtain the proper permission from the rights holder directly for these figures.

Export citation

Export citation

Choose format for download:

Download Citation
  • Open Access

Fisher Information Flow in Artificial Neural Networks

Maximilian Weimar1,*, Lukas M. Rachbauer1, Ilya Starshynov2, Daniele Faccio2, Linara Adilova3, Dorian Bouchet4, and Stefan Rotter1

  • *Contact author: maximilian.weimar@tuwien.ac.at

Phys. Rev. X 15, 031072 – Published 16 September, 2025

DOI: https://doi.org/10.1103/kn3z-rmm8

Abstract

The estimation of continuous parameters from measured data plays a central role in many fields of physics. A key tool in understanding and improving such estimation processes is the concept of Fisher information, which quantifies how information about unknown parameters propagates through a physical system and determines the ultimate limits of precision. With artificial neural networks gradually becoming an integral part of many measurement systems, it is essential to understand how they process and transmit parameter-relevant information internally. Here, we present a method to monitor the flow of Fisher information through an artificial neural network performing a parameter estimation task, tracking it from the input to the output layer. We show that optimal estimation performance corresponds to the maximal transmission of Fisher information and that training beyond this point results in information loss due to overfitting. This provides a model-free stopping criterion for network training—eliminating the need for a separate validation dataset. To demonstrate the practical relevance of our approach, we apply it to a network trained on data from an imaging experiment, highlighting its effectiveness in a realistic physical setting.

View figure in article

Physics Subject Headings (PhySH)

Popular Summary

Article Text

References (83)

  1. S. Yoon, M. Kim, M. Jang, Y. Choi, W. Choi, S. Kang, and W. Choi, Deep optical imaging within complex scattering media, Nat. Rev. Phys. 2, 141 (2020).
  2. J. Shlomi, P. Battaglia, and J.-R. Vlimant, Graph neural networks in particle physics, Mach. Learn. 2, 021001 (2020).
  3. T. Xie and J. C. Grossman, Crystal graph convolutional neural networks for an accurate and interpretable prediction of material properties, Phys. Rev. Lett. 120, 145301 (2018).
  4. M. Krenn, M. Malik, R. Fickler, R. Lapkiewicz, and A. Zeilinger, Automated search for new quantum experiments, Phys. Rev. Lett. 116, 090405 (2016).
  5. K. H. Jin, M. T. McCann, E. Froustey, and M. Unser, Deep convolutional neural network for inverse problems in imaging, IEEE Trans. Image Process. 26, 4509 (2017).
  6. T. M. Cover, Elements of Information Theory (John Wiley & Sons, Hoboken, New Jersey, 1999).
  7. D. J. MacKay, Information Theory, Inference and Learning Algorithms (Cambridge University Press, Cambridge, England, 2003).
  8. N. Tishby, F. C. Pereira, and W. Bialek, The information bottleneck method, arXiv:physics/0004057.
  9. B. C. Geiger and G. Kubin, Information bottleneck: Theory and applications in deep learning, Entropy 22, 1408 (2020).
  10. N. Tishby and N. Zaslavsky, Deep learning and the information bottleneck principle, in 2015 IEEE Information Theory Workshop (ITW) (IEEE, Jerusalem, Israel, 2015), pp. 1–5.
  11. Z. Goldfeld and Y. Polyanskiy, The information bottleneck problem and its applications in machine learning, IEEE J. Sel. Areas Inf. Theor. 1, 19 (2020).
  12. B. C. Geiger, On information plane analyses of neural network classifiers—a review, IEEE Trans. Neural Networks Learn. Syst. 33, 7039 (2021).
  13. A. Kolchinsky, B. D. Tracey, and D. H. Wolpert, Nonlinear information bottleneck, Entropy 21, 1181 (2019).
  14. T. Wu, I. Fischer, I. L. Chuang, and M. Tegmark, Learnability for the information bottleneck, in Proceedings of The 35th Uncertainty in Artificial Intelligence Conference, Proceedings of Machine Learning Research Vol. 115, edited by R. P. Adams and V. Gogate (PMLR, Tel Aviv, Israel, 2020), pp. 1050–1060.
  15. G. Neu, G. K. Dziugaite, M. Haghifam, and D. M. Roy, Information-theoretic generalization bounds for stochastic gradient descent, in Proceedings of Thirty Fourth Conference on Learning Theory, Proceedings of Machine Learning Research Vol. 134, edited by M. Belkin and S. Kpotufe (PMLR, Boulder, Colorado, 2021), pp. 3526–3545.
  16. H. Harutyunyan, M. Raginsky, G. Ver Steeg, and A. Galstyan, Information-theoretic generalization bounds for black-box learning algorithms, in Advances in Neural Information Processing Systems (Curran Associates, Inc., 2021), Vol. 34, pp. 24670–24682, https://proceedings.neurips.cc/paper_files/paper/2021/file/cf0d02ec99e61a64137b8a2c3b03e030-Paper.pdf.
  17. D. McAllester and K. Stratos, Formal limitations on the measurement of mutual information, in International Conference on Artificial Intelligence and Statistics (PMLR, Palermo, Italy, 2020), pp. 875–884.
  18. A. M. Saxe, Y. Bansal, J. Dapello, M. Advani, A. Kolchinsky, B. D. Tracey, and D. D. Cox, On the information bottleneck theory of deep learning*, J. Stat. Mech. (2019) 124020.
  19. Z. Goldfeld, E. van den Berg, K. Greenewald, I. Melnyk, N. Nguyen, B. Kingsbury, and Y. Polyanskiy, Estimating information flow in deep neural networks, arXiv:1810.05728.
  20. S. S. Lorenzen, C. Igel, and M. Nielsen, Information bottleneck: Exact analysis of (quantized) neural networks, in International Conference on Learning Representations (2021), arXiv:2106.12912.
  21. L. Adilova, B. C. Geiger, and A. Fischer, Information plane analysis for dropout neural networks, arXiv:2303.00596.
  22. A. Van den Bos, Parameter Estimation for Scientists and Engineers (John Wiley & Sons, Hoboken, New Jersey, 2007).
  23. P. Zegers, Fisher information properties, Entropy 17, 4918 (2015).
  24. H. H. Barrett, K. J. Myers, and S. Rathee, Foundations of Image Science (2004), ISBN [Amazon][WorldCat].
  25. Y. Shechtman, S. J. Sahl, A. S. Backer, and W. E. Moerner, Optimal point spread function design for 3D imaging, Phys. Rev. Lett. 113, 133902 (2014).
  26. Y. Li, Y. Xue, and L. Tian, Deep speckle correlation: A deep learning approach toward scalable imaging through scattering media, Optica 5, 1181 (2018).
  27. S.-L. Nyeo and R. R. Ansari, Data inversion for dynamic light scattering using Fisher information, Laser Phys. 25, 075703 (2015).
  28. C. Gonzalez-Ballestero, M. Aspelmeyer, L. Novotny, R. Quidant, and O. Romero-Isart, Levitodynamics: Levitation and control of microscopic objects in vacuum, Science 374, eabg3027 (2021).
  29. J. Hüpfl, F. Russo, L. M. Rachbauer, D. Bouchet, J. Lu, U. Kuhl, and S. Rotter, Continuity equation for the flow of Fisher information in wave scattering, Nat. Phys. 20, 1294 (2024).
  30. V. Giovannetti, S. Lloyd, and L. Maccone, Advances in quantum metrology, Nat. Photonics 5, 222 (2011).
  31. V. Gebhart, R. Santagati, A. A. Gentile, E. M. Gauger, D. Craig, N. Ares, L. Banchi, F. Marquardt, L. Pezzè, and C. Bonato, Learning quantum systems, Nat. Rev. Phys. 5, 141 (2023).
  32. T. Cokelaer, Parameter estimation of inspiralling compact binaries in ground-based detectors: Comparison between monte carlo simulations and the Fisher information matrix, Classical Quantum Gravity 25, 184007 (2008).
  33. M. Ježek, J. Fiurášek, and Z. c. v. Hradil, Quantum inference of states and processes, Phys. Rev. A 68, 012305 (2003).
  34. R. Pascanu and Y. Bengio, Revisiting natural gradient for deep networks, arXiv:1301.3584.
  35. R. Karakida, S. Akaho, and S.-i. Amari, Universal statistics of Fisher information in deep neural networks: Mean field approach, in Proceedings of the Twenty-Second International Conference on Artificial Intelligence and Statistics, Okinawa, Japan, edited by K. Chaudhuri and M. Sugiyama (PMLR, 2019), Vol. 89, pp. 1032–1041.
  36. J. Pennington and P. Worah, The spectrum of the Fisher information matrix of a single-hidden-layer neural network, in Advances in Neural Information Processing Systems, edited by S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (Curran Associates, Inc., Montr’eal, Canada, 2018), Vol. 31.
  37. J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska, D. Hassabis, C. Clopath, D. Kumaran, and R. Hadsell, Overcoming catastrophic forgetting in neural networks, Proc. Natl. Acad. Sci. U.S.A. 114, 3521 (2017).
  38. A. Achille, M. Rovere, and S. Soatto, Critical learning periods in deep neural networks, arXiv:1711.08856.
  39. Y. LeCun, Y. Bengio, and G. Hinton, Deep learning, Nature (London) 521, 436 (2015).
  40. F. Chollet, Deep Learning with Python (Simon and Schuster, Shelter Island, New York, 2021).
  41. R. G. Jarrett, Bounds and expansions for Fisher information when the moments are known, Biometrika 71, 101 (1984).
  42. P. Seriès, P. E. Latham, and A. Pouget, Tuning curve sharpening for orientation selectivity: Coding efficiency and the impact of correlations, Nat. Neurosci. 7, 1129 (2004).
  43. H. L. Van Trees, Detection, Estimation, and Modulation Theory, Part I: Detection, Estimation, and Linear Modulation Theory (John Wiley & Sons, New York, 2004).
  44. J. V. Beck and K. J. Arnold, Parameter Estimation in Engineering and Science (Wiley, New York, 1977), ISBN [Amazon][WorldCat].
  45. R. Zamir, A proof of the Fisher information inequality via a data processing argument, IEEE Trans. Inf. Theory 44, 1246 (1998).
  46. A. A. Alemi, I. Fischer, J. V. Dillon, and K. Murphy, Deep variational information bottleneck, arXiv:1612.00410.
  47. M. Stein, A. Mezghani, and J. A. Nossek, A lower bound for the Fisher information measure, IEEE Signal Process. Lett. 21, 796 (2014).
  48. B. Dulek, Characteristic function-based lower bounds for Fisher information under arbitrary parametrization, IEEE Trans. Aerospace Electron. Syst. 53, 501 (2017).
  49. M. S. Stein and J. A. Nossek, A pessimistic approximation for the Fisher information measure, IEEE Trans. Signal Process. 65, 386 (2016).
  50. M. S. Stein, M. Neumayer, and K. Barbé, Data-driven quality assessment of noisy nonlinear sensor and measurement systems, IEEE Trans. Instrum. Meas. 67, 1668 (2018).
  51. M. S. Stein, Sensitivity analysis for binary sampling systems via quantitative Fisher information lower bounds, IEEE Trans. Inf. Theory 68, 5601 (2022).
  52. D. Bouchet, S. Rotter, and A. P. Mosk, Maximum information states for coherent scattering measurements, Nat. Phys. 17, 564 (2021).
  53. I. Kanitscheider, R. Coen-Cagli, A. Kohn, and A. Pouget, Measuring Fisher information accurately in correlated neural populations, PLoS Comput. Biol. 11, e1004218 (2015).
  54. D. Bouchet, L. M. Rachbauer, S. Rotter, A. P. Mosk, and E. Bossy, Optimal control of coherent light scattering for binary decision problems, Phys. Rev. Lett. 127, 253902 (2021).
  55. E. Z. Gross, Information loss in a non-linear neuronal model, Physica (Amsterdam) 525A, 443 (2019).
  56. F. Rosenblatt, The perceptron: A probabilistic model for information storage and organization in the brain, Psychol. Rev. 65, 386 (1958).
  57. G.-B. Huang, What are extreme learning machines? filling the gap between Frank Rosenblatt’s dream and John von Neumann’s puzzle, Cognitive Comput. 7, 263 (2015).
  58. J. Alsing and B. Wandelt, Generalized massive optimal data compression, Mon. Not. R. Astron. Soc. 476, L60 (2018).
  59. W. Liang, M. Manry, Q. Yu, S. Apollo, M. Dawson, and A. Fung, Bounding the performance of neural network estimators, given only a set of training data, in Proceedings of 1994 28th Asilomar Conference on Signals, Systems and Computers (IEEE, Pacific Grove, CA, 1994), Vol. 2, pp. 912–916.
  60. M. T. Manry, C.-H. Hsieh, M. S. Dawson, A. K. Fung, and S. J. Apollo, Cramer Rao maximum a-posteriori bounds on neural network training error for non-Gaussian signals and parameters (1996), pp. 381–391.
  61. M. T. Manry, S. J. Apollo, and Q. Yu, Minimum mean square estimation and neural networks, Neurocomputing;Variable Star Bulletin 13, 59 (1996).
  62. X. Zhang, Q. Duchemin, K. Liu*, C. Gultekin, S. Flassbeck, C. Fernandez-Granda, and J. Assländer, Cramér–rao bound-informed training of neural networks for quantitative mri, Magn. Reson. Med. 88, 436 (2022).
  63. S. P. Nolan, L. Pezzè, and A. Smerzi, Frequentist parameter estimation with supervised learning, AVS Quantum Sci. 3, 034401 (2021).
  64. J. Lalis, B. Gerardo, and Y. Byun, An adaptive stopping criterion for backpropagation learning in feedforward neural network, Int. J. Multimedia Ubiquitous Eng. 9, 149 (2014).
  65. L. Prechelt, Early stopping-but when?, in Neural Networks: Tricks of the Trade (Springer, New York, 2002), pp. 55–69.
  66. T. Miseta, A. Fodor, and Á. Vathy-Fogarassy, Surpassing early stopping: A novel correlation-based stopping criterion for neural networks, Neurocomputing;Variable Star Bulletin 567, 127028 (2024).
  67. S. Natarajan and R. R. Rhinehart, Automated stopping criteria for neural network training, in Proceedings of the 1997 American Control Conference (Cat. No. 97CH36041) (IEEE, New York, 1997), Vol. 4, pp. 2409–2413.
  68. M. V. Ferro, Y. D. Mosquera, F. J. R. Pena, and V. M. D. Bilbao, Early stopping by correlating online indicators in neural networks, Neural Netw. 159, 109 (2023).
  69. C. M. Ennett, M. Frize, and N. Scales, Evaluation of the logarithmic-sensitivity index as a neural network stopping criterion for rare outcomes, in Proceedings of the 4th International IEEE EMBS Special Topic Conference on Information Technology Applications in Biomedicine, 2003. (IEEE, New York, 2003), pp. 207–210.
  70. V. Berisha and A. O. Hero, Empirical non-parametric estimation of the Fisher information, IEEE Signal Process. Lett. 22, 988 (2014).
  71. T. T. Duy, L. V. Nguyen, V.-D. Nguyen, N. L. Trung, and K. Abed-Meraim, Fisher information neural estimation, in 2022 30th European Signal Processing Conference (EUSIPCO) (IEEE, Belgrade, Serbia, 2022), pp. 2111–2115.
  72. L. Livi, F. M. Bianchi, and C. Alippi, Determination of the edge of criticality in echo state networks through Fisher information maximization, IEEE Trans. Neural Networks Learn. Syst. 29, 706 (2017).
  73. O. Har-Shemesh, R. Quax, B. Miñano, A. G. Hoekstra, and P. M. A. Sloot, Nonparametric estimation of Fisher information from real data, Phys. Rev. E 93, 023301 (2016).
  74. K. He, X. Zhang, S. Ren, and J. Sun, Deep residual learning for image recognition, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (IEEE, New York, 2016), pp. 770–778.
  75. L. G. Wright, T. Onodera, M. M. Stein, T. Wang, D. T. Schachter, Z. Hu, and P. L. McMahon, Deep physical neural networks trained with backpropagation, Nature (London) 601, 549 (2022).
  76. A. Momeni, B. Rahmani, B. Scellier, L. G. Wright, P. L. McMahon, C. C. Wanjura, Y. Li, A. Skalli, N. G. Berloff, T. Onodera et al., Training of physical neural networks, arXiv:2406.03372.
  77. C. C. Wanjura and F. Marquardt, Fully nonlinear neuromorphic computing with linear wave scattering, Nat. Phys. 20, 1434 (2024).
  78. M. Weimar, Experimental data set for calculating the Fisher information in an artificial neural network, 10.5281/zenodo.17021435.
  79. M. Weimar, Calculation of Fisher information inside the layers of an artificial neural network that performs a parameter-estimation task, https://github.com/maximilianweimar/Fisher-information-flow-in-artificial-neural-networks.
  80. D. Pollard, A note on insufficiency and the preservation of Fisher information, in From Probability to Statistics and Back: High-Dimensional Models and Processes–A Festschrift in Honor of Jon A. Wellner (Institute of Mathematical Statistics, Beachwood, OH, 2013), Vol. 9, pp. 266–276.
  81. A. Serbes and K. A. Qaraqe, Threshold regions in frequency estimation, IEEE Trans. Aerospace Electron. Syst. 58, 4850 (2022).
  82. A. Papoulis and J. G. Hoffman, Probability, random variables, and stochastic processes, Phys. Today 20, No. 1, 135 (1967).
  83. A. Andersson, Mechanisms for log normal concentration distributions in the environment, Sci. Rep. 11, 16418 (2021).

Outline

Information

Sign In to Your Journals Account

Filter

Filter

Article Lookup

Enter a citation