- Open Access
Bilinear Sequence Regression: A Model for Learning from Long Sequences of High-Dimensional Tokens
Phys. Rev. X 15, 021092 – Published 16 June, 2025
DOI: https://doi.org/10.1103/l4p2-vrxt
Abstract
Current progress in artificial intelligence is centered around so-called large language models that consist of neural networks processing long sequences of high-dimensional vectors called tokens. Statistical physics provides powerful tools to study the functioning of learning with neural networks and has played a recognized role in the development of modern machine learning. The statistical physics approach relies on simplified and analytically tractable models of data. However, simple tractable models for long sequences of high-dimensional tokens are largely underexplored. Inspired by the crucial role models such as the single-layer teacher-student perceptron (also known as generalized linear regression) played in the theory of fully connected neural networks, in this paper, we introduce and study the bilinear sequence regression (BSR) as one of the most basic models for sequences of tokens. We note that modern architectures naturally subsume the BSR model due to the skip connections. Building on recent methodological progress, we compute the Bayes-optimal generalization error for the model in the limit of long sequences of high-dimensional tokens and provide a message-passing algorithm that matches this performance. We quantify the improvement that optimal learning brings with respect to vectorizing the sequence of tokens and learning via simple linear regression. We also unveil surprising properties of the gradient descent algorithms in the BSR model.
Physics Subject Headings (PhySH)
Popular Summary
Modern AI systems have achieved remarkable results by learning from long sequences of complex inputs like natural language or biological data. Yet, the theoretical understanding of how neural networks learn from such structured data remains limited. In this work, we introduce a simplified but powerful model to study this learning process from first principles. The model—bilinear sequence regression (BSR)—captures some of the core features of advanced architectures like transformers while remaining analytically solvable.
The BSR model defines outputs as structured functions of sequences of high-dimensional vectors, enabling precise analysis of learning behavior. Using techniques from statistical physics and probabilistic inference, we determine the optimal prediction accuracy in the “thermodynamic limit,” where both the sequence length and token dimension become large. This analysis uncovers a sharp phase transition in learning: There exists a threshold in data volume below which learning fails and above which it suddenly becomes successful. We also analyze gradient descent learning dynamics and provide numerical evidence that a variant of this algorithm can reach the optimal performance predicted by our theory.
BSR provides a rigorous and flexible framework for understanding how neural networks learn from sequential data. It offers insight into why architectures like transformers perform so well and helps to identify the fundamental conditions under which they succeed or fail.
Article Text
References (88)
- Y. LeCun, Y. Bengio, and G. Hinton, Deep learning, Nature (London) 521, 436 (2015).
- J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, ImageNet: A large-scale hierarchical image database, in Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition (IEEE, New York, 2009), pp. 248–255.
- A. Krizhevsky, I. Sutskever, and G. E. Hinton, Imagenet classification with deep convolutional neural networks, Adv. Neural Inf. Process. Syst. 25 (2012).
- D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot et al., Mastering the game of go with deep neural networks and tree search, Nature (London) 529, 484 (2016).
- L. Zdeborová, Understanding deep learning is also a job for physicists, Nat. Phys. 16, 602 (2020).
- J. J. Hopfield, Neural networks and physical systems with emergent collective computational abilities, Proc. Natl. Acad. Sci. U.S.A. 79, 2554 (1982).
- D. H. Ackley, G. E. Hinton, and T. J. Sejnowski, A learning algorithm for Boltzmann machines, Cogn. Sci. 9, 147 (1985).
- E. Gardner and B. Derrida, Optimal storage properties of neural network models, J. Phys. A 21, 271 (1988).
- E. Gardner and B. Derrida, Three unfinished works on the optimal storage capacity of networks, J. Phys. A 22, 1983 (1989).
- H. S. Seung, H. Sompolinsky, and N. Tishby, Statistical mechanics of learning from examples, Phys. Rev. A 45, 6056 (1992).
- A. M. Saxe, J. L. McClelland, and S. Ganguli, Exact solutions to the nonlinear dynamics of learning in deep linear neural networks, International Conference on Learning Representations (2014).
- C. Baldassi, A. Ingrosso, C. Lucibello, L. Saglietti, and R. Zecchina, Subdominant dense clusters allow for simple learning and high computational performance in neural networks with discrete synapses, Phys. Rev. Lett. 115, 128101 (2015).
- C. Baldassi, C. Borgs, J. T. Chayes, A. Ingrosso, C. Lucibello, L. Saglietti, and R. Zecchina, Unreasonable effectiveness of learning neural networks: From accessible states and robust ensembles to basic algorithmic schemes, Proc. Natl. Acad. Sci. U.S.A. 113, E7655 (2016).
- J. Barbier, F. Krzakala, N. Macris, L. Miolane, and L. Zdeborová, Optimal errors and phase transitions in high-dimensional generalized linear models, Proc. Natl. Acad. Sci. U.S.A. 116, 5451 (2019).
- S. Goldt, M. Mézard, F. Krzakala, and L. Zdeborová, Modeling the influence of data structure on learning in neural networks: The hidden manifold model, Phys. Rev. X 10, 041044 (2020).
- M. S. Advani and A. M. Saxe, High-dimensional dynamics of generalization error in neural networks, Neural Netw. 132, 428 (2020).
- B. Loureiro, C. Gerbelot, H. Cui, S. Goldt, F. Krzakala, M. Mezard, and L. Zdeborová, Learning curves of generic features maps for realistic datasets with a teacher-student model, Adv. Neural Inf. Process. Syst. 34, 18137 (2021).
- B. Sorscher, R. Geirhos, S. Shekhar, S. Ganguli, and A. Morcos, Beyond neural scaling laws: Beating power law scaling via data pruning, Adv. Neural Inf. Process. Syst. 35, 19523 (2022).
- A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, Attention is all you need, Adv. Neural Inf. Process. Syst. 30 (2017).
- A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al., Language models are unsupervised multitask learners, OpenAI blog 1, 9 (2019).
- B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal et al., Language models are few-shot learners, arXiv:2005.14165.
- J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat et al., GPT-4 technical report, arXiv:2303.08774.
- S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg et al., Sparks of artificial general intelligence: Early experiments with GPT-4, arXiv:2303.12712.
- J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, E. H. Chi, T. Hashimoto, O. Vinyals, P. Liang, J. Dean, and W. Fedus, Emergent abilities of large language models, Trans. Mach. Learn. Res. (2022).
- J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, Scaling laws for neural language models, arXiv:2001.08361.
- T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al., Language models are few-shot learners, Adv. Neural Inf. Process. Syst. 33, 1877 (2020).
- G. Cybenko, Approximation by superpositions of a sigmoidal function, Math. Control Signals Syst. 2, 303 (1989).
- K. Hornik, M. Stinchcombe, and H. White, Multilayer feedforward networks are universal approximators, Neural Netw. 2, 359 (1989).
- A. Raventós, M. Paul, F. Chen, and S. Ganguli, Pretraining task diversity and the emergence of non-Bayesian in-context learning for regression, Adv. Neural Inf. Process. Syst. 36, 14228 (2023).
- F. Cagnetta and M. Wyart, Towards a theory of how the structure of language is acquired by deep neural networks, arXiv:2406.00048.
- F. Behrens, L. Biggio, and L. Zdeborová, Understanding counting in small transformers: The interplay between attention and feed-forward layers, arXiv:2407.11542.
- B. Geshkovski, C. Letrouit, Y. Polyanskiy, and P. Rigollet, A mathematical perspective on transformers, arXiv:2312.10794.
- A. Cowsik, T. Nebabu, X.-L. Qi, and S. Ganguli, Geometric dynamics of signal propagation predict trainability of transformers, arXiv:2403.02579.
- R. Rende, F. Gerace, A. Laio, and S. Goldt, Mapping of attention mechanisms to a generalized Potts model, Phys. Rev. Res. 6, 023057 (2024).
- H. Cui, F. Behrens, F. Krzakala, and L. Zdeborová, A phase transition between positional and semantic learning in a solvable model of dot-product attention, Adv. Neural Inf. Process. Syst. 37, 36342 (2024).
- Y. M. Lu, M. I. Letey, J. A. Zavatone-Veth, A. Maiti, and C. Pehlevan, Asymptotic theory of in-context learning by linear attention, arXiv:2405.11751.
- A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan et al., The Llama 3 herd of models, arXiv:2407.21783.
- F. Mignacco, P. Urbani, and L. Zdeborová, Stochasticity helps to navigate rough landscapes: Comparing gradient-descent-based algorithms in the phase retrieval problem, Mach. Learn. 2, 035029 (2021).
- S. Sarao Mannelli, G. Biroli, C. Cammarota, F. Krzakala, P. Urbani, and L. Zdeborová, Complex dynamics in simple neural networks: Understanding gradient flow in phase retrieval, Adv. Neural Inf. Process. Syst. 33, 3265 (2020).
- M. Belkin, D. Hsu, S. Ma, and S. Mandal, Reconciling modern machine-learning practice and the classical bias–variance trade-off, Proc. Natl. Acad. Sci. U.S.A. 116, 15849 (2019).
- F. Gerace, B. Loureiro, F. Krzakala, M. Mézard, and L. Zdeborová, Generalisation error in learning with random features and the hidden manifold model, in Proceedings of the 37th International Conference on Machine Learning, Proceedings of Machine Learning Research Vol. 119 (PMLR, Cambridge, MA, 2020), pp. 3452–3462.
- C. Schülke, P. Schniter, and L. Zdeborová, Phase diagram of matrix compressed sensing, Phys. Rev. E 94, 062136 (2016).
- https://github.com/SPOC-group/bilinear-sequence-regression.
- I. Tolstikhin et al., Mlp-mixer: An all-mlp architecture for vision, Adv. Neural Inf. Process. Syst. 34, 24261 (2021).
- B. Recht, M. Fazel, and P. A. Parrilo, Guaranteed minimum-rank solutions of linear matrix equations via nuclear norm minimization, SIAM Rev. 52, 471 (2010).
- D. L. Donoho, M. Gavish, and A. Montanari, The phase transition of matrix recovery from Gaussian measurements matches the minimax MSE of matrix denoising, Proc. Natl. Acad. Sci. U.S.A. 110, 8405 (2013).
- C. Giraud, Low rank multivariate regression, Electron. J. Stat. 5, 775 (2011).
- P. D. Hoff, Multilinear tensor regression for longitudinal relational data, Ann. Appl. Stat. 9, 1169 (2015).
- E. Y. Chen and J. Fan, Statistical inference for high-dimensional matrix-variate factor models, J. Am. Stat. Assoc. 118, 1038 (2023).
- S. Gunasekar, B. E. Woodworth, S. Bhojanapalli, B. Neyshabur, and N. Srebro, Implicit regularization in matrix factorization, Adv. Neural Inf. Process. Syst. 30 (2017).
- Z. Li, Y. Luo, and K. Lyu, Towards resolving the implicit bias of gradient descent for matrix factorization: Greedy low-rank learning, in Proceedings of the International Conference on Learning Representations (2021).
- H. C. Schmidt, Statistical physics of sparse and dense models in optimization and inference, Ph.D. thesis, Université Paris Saclay (COmUE), 2018.
- A. Maillard, F. Krzakala, M. Mézard, and L. Zdeborová, Perturbative construction of mean-field equations in extensive-rank matrix factorization and denoising, J. Stat. Mech. (2022) 083301.
- J. Barbier and N. Macris, Statistical limits of dictionary learning: Random matrix theory and the spectral replica method, Phys. Rev. E 106, 024136 (2022).
- E. Troiani, V. Erba, F. Krzakala, A. Maillard, and L. Zdeborová, Optimal denoising of rotationally invariant rectangular matrices, in Mathematical and Scientific Machine Learning, Proceedings of Machine Learning Research Vol. 190, edited by B. Dong, Q. Li, L. Wang, and Z.-Q. J. Xu (PMLR, Cambridge, MA, 2022), pp. 97–112.
- G. Semerjian, Matrix denoising: Bayes-optimal estimators via low-degree polynomials, J. Stat. Phys. 191, 139 (2024).
- F. Camilli and M. Mézard, Matrix factorization with neural networks, Phys. Rev. E 107, 064308 (2023).
- F. Pourkamali and N. Macris, Rectangular rotational invariant estimator for general additive noise matrices, in Proceedings of the 2023 IEEE International Symposium on Information Theory (ISIT) (IEEE, New York, 2023), pp. 2081–2086.
- Y. V. Fyodorov, A spin glass model for reconstructing nonlinearly encrypted signals corrupted by noise, J. Stat. Phys. 175, 789 (2019).
- P. J. Kamali and P. Urbani, Dynamical mean field theory for models of confluent tissues and beyond, SciPost Phys. 15, 219 (2023).
- A. Montanari and E. Subag, Solving overparametrized systems of random equations: I. Model and algorithms for approximate solutions, arXiv:2306.13326.
- P. J. Kamali and P. Urbani, Stochastic gradient descent outperforms gradient descent in recovering a high-dimensional signal in a glassy energy landscape, arXiv:2309.04788.
- A. Maillard, E. Troiani, S. Martin, F. Krzakala, and L. Zdeborová, Bayes-optimal learning of an extensive-width neural network from quadratically many samples, in Advances in Neural Information Processing Systems (Curran Associates, Inc., 2024), Vol. 37, pp. 82085–82132, https://proceedings.neurips.cc/paper_files/paper/2024/file/953e742190ca02fc8f9f710052f2fead-Paper-Conference.pdf.
- S. Bhojanapalli, B. Neyshabur, and N. Srebro, Global optimality of local search for low rank matrix recovery, Adv. Neural Inf. Process. Syst. 29 (2016).
- L. Ding, D. Drusvyatskiy, M. Fazel, and Z. Harchaoui, Flat minima generalize for low-rank matrix recovery, Inf. Inference 13, iaae009 (2024).
- K. Okajima and T. Takahashi, Asymptotic dynamics of alternating minimization for bilinear regression, arXiv:2402.04751.
- H. Hu and Y. M. Lu, Universality laws for high-dimensional learning with random features, IEEE Trans. Inf. Theory 69, 1932 (2022).
- Z. Wang, E. Nichani, and J. D. Lee, Learning hierarchical polynomials with three-layer neural networks, in Proceedings of the Twelfth International Conference on Learning Representations (2024).
- T. M. Cover and J. A. Thomas, Information theory and statistics, Elem. Inf. Theor. 1, 279 (1991).
- L. Zdeborová and F. Krzakala, Statistical physics of inference: Thresholds and algorithms, Adv. Phys. 65, 453 (2016).
- J. Pennington and P. Worah, Nonlinear random matrix theory for deep learning, Adv. Neural Inf. Process. Syst. 30 (2017).
- P. Biane, On the free convolution with a semi-circular distribution, Indiana Univ. Math. J. 46, 705 (1997).
- S. Rangan, Generalized approximate message passing for estimation with random linear mixing, in Proceedings of the 2011 IEEE International Symposium on Information Theory Proceedings (IEEE, New York, 2011), pp. 2168–2172.
- D. L. Donoho, A. Maleki, and A. Montanari, Message-passing algorithms for compressed sensing, Proc. Natl. Acad. Sci. U.S.A. 106, 18914 (2009).
- R. Berthier, A. Montanari, and P.-M. Nguyen, State evolution for approximate message passing with non-separable functions, Inf. Inference 9, 33 (2020).
- C. Gerbelot and R. Berthier, Graph-based approximate message passing iterations, Inf. Inference 12, 2562 (2023).
- E. Romanov and M. Gavish, Near-optimal matrix recovery from random linear measurements, Proc. Natl. Acad. Sci. U.S.A. 115, 7200 (2018).
- J. Barbier, J. Ko, and A. A. Rahman, Information-theoretic limits for sublinear-rank symmetric matrix factorization, in Proceedings of the International Zurich Seminar on Information and Communication (IZS 2024) (ETH Zürich, 2024), p. 16.
- J. Barbier, J. Ko, and A. A. Rahman, A multiscale cavity method for sublinear-rank symmetric matrix factorization, arXiv:2403.07189.
- F. Pourkamali, J. Barbier, and N. Macris, Matrix inference in growing rank regimes, IEEE Trans. Inf. Theory (to be published).
- U. Helmke and J. B. Moore, Optimization and Dynamical Systems (Springer Science & Business Media, New York, 2012).
- D. Donoho and M. Gavish, Minimax risk of matrix denoising by singular value thresholding, Ann. Stat. 42, 2413 (2014).
- F. Krzakala, M. Mézard, F. Sausset, Y. F. Sun, and L. Zdeborová, Statistical-physics-based reconstruction in compressed sensing, Phys. Rev. X 2, 021005 (2012).
- B. Neyshabur, R. Tomioka, and N. Srebro, In search of the real inductive bias: On the role of implicit regularization in deep learning, arXiv:1412.6614.
- S. Arora, N. Cohen, W. Hu, and Y. Luo, Implicit regularization in deep matrix factorization, Adv. Neural Inf. Process. Syst. 32 (2019).
- F. Mignacco, F. Krzakala, P. Urbani, and L. Zdeborová, Dynamical mean-field theory for stochastic gradient descent in Gaussian mixture classification, Adv. Neural Inf. Process. Syst. 33, 9540 (2020).
- S. Diamond and S. Boyd, CVXPY: A Python-embedded modeling language for convex optimization, J. Mach. Learn. Res. 17, 1 (2016).
- S. D. Akshay Agrawal, Robin Verschueren, and S. Boyd, A rewriting system for convex optimization problems, J. Control Decis. 5, 42 (2018).
