- Open Access
Statistical Physics Analysis of Graph Neural Networks: Approaching Optimality in the Contextual Stochastic Block Model
Phys. Rev. X 15, 041026 – Published 10 November, 2025
DOI: https://doi.org/10.1103/lfxj-hbsk
Abstract
Graph neural networks (GNNs) are designed to process data associated with graphs. They are finding an increasing range of applications; however, as with other modern machine learning techniques, their theoretical understanding is limited. GNNs can encounter difficulties in gathering information from nodes that are far apart by iterated aggregation steps. This situation is partly caused by so-called oversmoothing; and overcoming it is one of the practically motivated challenges. We consider the situation where information is aggregated by multiple steps of convolution, leading to graph convolutional networks (GCNs). We analyze the generalization performance of a basic GCN, trained for node classification on data generated by the contextual stochastic block model. We predict its asymptotic performance by deriving the free energy of the problem, using the replica method, in the high-dimensional limit. Calling depth the number of convolutional steps, we show the importance of going to large depth to approach the Bayes optimality. We detail how the architecture of the GCN has to scale with the depth to avoid oversmoothing. The resulting large depth limit can be close to the Bayes optimality and leads to a continuous GCN. Technically, we tackle this continuous limit via an approach that resembles dynamical mean-field theory with constraints at the initial and final times. An expansion around large regularization allows us to solve the corresponding equations for the performance of the deep GCN. This promising tool may contribute to the analysis of further deep neural networks.
Physics Subject Headings (PhySH)
Popular Summary
Graphs are a fundamental way to represent complex systems, from molecules to social networks to interactions in the brain. Graph neural networks (GNNs) have become a popular tool to analyze such data, but the theoretical understanding of these networks lags behind their success. A central open question is whether making GNNs deeper—so that they can combine information from more distant nodes—actually improves performance without losing crucial details. In this study, we provide a precise answer by analyzing how depth affects accuracy in graph convolutional networks (GCNs), a simple but representative form of GNN.
We examine GCNs on a standard benchmark in which nodes belong to hidden classes, tend to connect more often with nodes of the same class, and carry features that offer additional clues. The task is to infer the classes of most nodes using both the graph structure and node features, given only a few known labels. By applying tools from statistical physics, we calculate the exact performance of the GCN in the limit of very large graphs. Our results show that deeper networks are indeed necessary, but they must be carefully designed: Residual connections and strong regularization allow GCNs with many layers to approach the theoretical optimal performance.
This work delivers the first exact predictions for the generalization ability of infinitely deep, or continuous, GNNs. It provides clear guidelines for avoiding oversmoothing and explains when depth is beneficial. Looking ahead, the same physics-based approach can be applied to more advanced models, such as attention-based networks, bringing us closer to a systematic theory of deep learning on graphs.
Article Text
References (57)
- Y. Wang, Z. Li, and A. Barati Farimani, Graph neural networks for molecules, in Machine Learning in Molecular Sciences (Springer International Publishing, Cham, 2023), pp. 21–66, .
- M. M. Li, K. Huang, and M. Zitnik, Graph representation learning in biomedicine and healthcare, Nat. Biomed. Eng. 6, 1353 (2022).
- A. Bessadok, M. A. Mahjoub, and I. Rekik, Graph neural networks in network neuroscience, arXiv:2106.03535.
- A. Sanchez-Gonzalez, J. Godwin, T. Pfaff, R. Ying, J. Leskovec, and P. W. Battaglia, Learning to simulate complex physics with graph networks, in Proceedings of the 37th International Conference on Machine Learning (2020), arXiv:2002.09405.
- J. Shlomi, P. Battaglia, and J.-R. Vlimant, Graph neural networks in particle physics, Mach. Learn. 2, 021001 (2020).
- Y. Peng, B. Choi, and J. Xu, Graph learning for combinatorial optimization: A survey of state-of-the-art, Data Sci. Eng. 6, 119 (2021).
- Q. Cappart, D. Chételat, E. Khalil, A. Lodi, C. Morris, and P. Veličković, Combinatorial optimization and reasoning with graph neural networks, J. Mach. Learn. Res. 24, 1 (2023).
- C. Morris, F. Frasca, N. Dym, H. Maron, I. I. Ceylan, R. Levie, D. Lim, M. Bronstein, M. Grohe, and S. Jegelka, Position: Future directions in the theory of graph machine learning, in Proceedings of the 41st International Conference on Machine Learning (2024).
- Q. Li, Z. Han, and X.-M. Wu, Deeper insights into graph convolutional networks for semi-supervised learning, in Thirty-Second AAAI Conference on Artificial Intelligence (2018) arXiv:1801.07606.
- K. Oono and T. Suzuki, Graph neural networks exponentially lose expressive power for node classification, in International Conference on Learning Representations (2020), arXiv:1905.10947.
- G. Li, M. Müller, A. Thabet, and B. Ghanem, DeepGCNs: Can GCNs go as deep as CNNs?, in ICCV (2019), arXiv:1904.03751.
- M. Chen, Z. Wei, Z. Huang, B. Ding, and Y. Li, Simple and deep graph convolutional networks, in Proceedings of the 37th International Conference on Machine Learning (2020), arXiv:2007.02133.
- H. Ju, D. Li, A. Sharma, and H. R. Zhang, Generalization in graph neural networks: Improved PAC-Bayesian bounds on graph diffusion, in AISTATS (2023), arXiv:2302.04451.
- H. Tang and Y. Liu, Towards understanding the generalization of graph neural networks (2023), arXiv:2305.08048.
- W. Cong, M. Ramezani, and M. Mahdavi, On provable benefits of depth in training graph convolutional networks, in 35th Conference on Neural Information Processing Systems (2021), arXiv:2110.15174.
- P. M. Esser, L. C. Vankadara, and D. Ghoshdastidar, Learning theory can (sometimes) explain generalisation in graph neural networks, in 35th Conference on Neural Information Processing Systems (2021), arXiv:2112.03968.
- H. S. Seung, H. Sompolinsky, and N. Tishby, Statistical mechanics of learning from examples, Phys. Rev. A 45, 6056 (1992).
- B. Loureiro, C. Gerbelot, H. Cui, S. Goldt, F. Krzakala, M. Mezard, and L. Zdeborová, Learning curves of generic features maps for realistic datasets with a teacher-student model, Adv. Neural Inf. Process. Syst. 34, 18137 (2021).
- S. Mei and A. Montanari, The generalization error of random features regression: Precise asymptotics and the double descent curve, Commun. Pure Appl. Math. 75, 667 (2022).
- C. Shi, L. Pan, H. Hu, and I. Dokmanić, Homophily modulates double descent generalization in graph convolution networks, Proc. Natl. Acad. Sci. U.S.A. 121, e2309504121 (2023).
- B. Yan and P. Sarkar, Covariate regularized community detection in sparse graphs, J. Am. Stat. Assoc. 116, 734 (2021).
- Y. Deshpande, S. Sen, A. Montanari, and E. Mossel, Contextual stochastic block models, in Advances in Neural Information Processing Systems, edited by S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (2018), Vol. 31, .
- E. Chien, J. Peng, P. Li, and O. Milenkovic, Adaptative universal generalized pagerank graph neural network, in Proceedings of the 39th International Conference on Learning Representations (2021), arXiv:2006.07988.
- G. Fu, P. Zhao, and Y. Bian, p-Laplacian based graph neural networks, in Proceedings of the 39th International Conference on Machine Learning (2021), arXiv:2111.07337.
- R. Lei, Z. Wang, Y. Li, B. Ding, and Z. Wei, EvenNet: Ignoring odd-hop neighbors improves robustness of graph neural networks, in 36th Conference on Neural Information Processing Systems (2022), arXiv:2205.13892.
- O. Duranthon and L. Zdeborová, Asymptotic generalization error of a single-layer graph convolutional network, in The Learning on Graphs Conference (2024), arXiv:2402.03818.
- R. T. Q. Chen, Y. Rubanova, J. Bettencourt, and D. Duvenaud, Neural ordinary differential equations, in 32nd Conference on Neural Information Processing Systems (2018), arXiv:1806.07366.
- T. N. Kipf and M. Welling, Semi-supervised classification with graph convolutional networks, in International Conference on Learning Representations (2017), arXiv:1609.02907.
- H. Cui, F. Krzakala, and L. Zdeborová, Bayes-optimal learning of deep random networks of extensive-width, in Proceedings of the 40th International Conference on Machine Learning (2023), arXiv:2302.00375.
- A. K. McCallum, K. Nigam, J. Rennie, and K. Seymore, Automating the construction of internet portals with machine learning, Information Retrieval 3, 127 (2000).
- O. Shchur, M. Mumme, A. Bojchevski, and S. Günnemann, Pitfalls of graph neural network evaluation, arXiv:1811.05868.
- C. L. Giles, K. D. Bollacker, and S. Lawrenc, Citeseer: An automatic citation indexing system, in Proceedings of the Third ACM Conference on Digital Libraries (1998), pp. 89–98.
- P. Sen, G. Namata, M. Bilgic, L. Getoor, B. Galligher, and T. Eliassi-Rad, Collective classification in network data, AI Magazine 29, 93 (2008).
- A. Baranwal, K. Fountoulakis, and A. Jagannath, Graph convolution for semi-supervised classification: Improved linear separability and out-of-distribution generalization, in Proceedings of the 38th International Conference on Machine Learning (2021), arXiv:2102.06966.
- A. Baranwal, K. Fountoulakis, and A. Jagannath, Optimality of message-passing architectures for sparse graphs, in 37th Conference on Neural Information Processing Systems (2023), arXiv:2305.10391.
- R. Wang, A. Baranwal, and K. Fountoulakis, Analysis of corrected graph convolutions, arXiv:2405.13987.
- F. Mignacco, F. Krzakala, Y. M. Lu, and L. Zdeborová, The role of regularization in classification of high-dimensional noisy Gaussian mixture, in International Conference on Learning Representations (2020), arXiv:2002.11544.
- B. Aubin, F. Krzakala, Y. M. Lu, and L. Zdeborová, Generalization error in high-dimensional perceptrons: Approaching Bayes error with convex optimization, in Advances in Neural Information Processing Systems (2020), arXiv:2006.06560.
- O. Duranthon and L. Zdeborová, Optimal inference in contextual stochastic block models, arXiv:2306.07948 [Trans. Mach. Learn. Res. (to be published)], https://openreview.net/forum?id=Pe6hldOUkw.
- N. Keriven, Not too little, not too much: A theoretical analysis of graph (over)smoothing, in 36th Conference on Neural Information Processing Systems (2022), arXiv:2205.12156.
- K. He, X. Zhang, S. Ren, and J. Sun, Deep residual learning for image recognition, in IEEE Conference on Computer Vision and Pattern Recognition (2016), arXiv:1512.03385.
- T. Pham, T. Tran, D. Phung, and S. Venkatesh, Column networks for collective classification, in AAAI (2017), .
- K. Xu, M. Zhang, S. Jegelka, and K. Kawaguchi, Optimization of graph neural networks: Implicit acceleration by skip connections and more depth, in Proceedings of the 38th International Conference on Machine Learning (2021), arXiv:2105.04550.
- M. E. Sander, P. Ablin, and G. Peyré, Do residual neural networks discretize neural ordinary differential equations?, in 36th Conference on Neural Information Processing Systems (2022), arXiv:2205.14612.
- J. Ling, A. Kurzawski, and J. Templeton, Reynolds averaged turbulence modelling using deep neural networks with embedded invariance, J. Fluid Mech. 807, 155 (2016).
- C. Rackauckas, Y. Ma, J. Martensen, C. Warner, K. Zubov, R. Supekar, D. Skinner, A. Ramadhan, and A. Edelman, Universal differential equations for scientific machine learning, arXiv:2001.04385.
- P. Marion, Generalization bounds for neural ordinary differential equations and deep residual networks, arXiv:2305.06648.
- M. Poli, S. Massaroli, J. Park, A. Yamashita, H. Asama, and J. Park, Graph neural ordinary differential equations, arXiv:1911.07532.
- L.-P. A. C. Xhonneux, M. Qu, and J. Tang, Continuous graph neural networks, in Proceedings of the 37th International Conference on Machine Learning (2020), arXiv:1912.00967.
- A. Han, D. Shi, L. Lin, and J. Gao, From continuous dynamics to graph neural networks: Neural diffusion and beyond, arXiv:2310.10121.
- O. Duranthon and L. Zdeborová (2025), https://arxiv.org/src/2503.01361v2/anc.
- C. Lu and S. Sen, Contextual stochastic block model: Sharp thresholds and contiguity, arXiv:2011.09841.
- F. Wu, T. Zhang, A. H. de Souza Jr., C. Fifty, T. Yu, and K. Q. Weinberger, Simplifying graph convolutional networks, in Proceedings of the 36th International Conference on Machine Learning (2019), arXiv:1902.07153.
- H. Zhu and P. Koniusz, Simple spectral graph convolution, in International Conference on Learning Representations (2021).
- T. Lesieur, F. Krzakala, and L. Zdeborová, Constrained low-rank matrix estimation: Phase transitions, approximate message passing and applications, J. Stat. Mech. (2017) 073403.
- O. Duranthon and L. Zdeborová, Neural-prior stochastic block model, Mach. Learn. 4, 035017 (2023).
- J. Baik, G. B. Arous, and S. Péché, Phase transition of the largest eigenvalue for nonnull complex sample covariance matrices, Ann. Probab. 33, 1643 (2005).
