Reuse & Permissions

It is not necessary to obtain permission to reuse this article or its components as it is available under the terms of the Creative Commons Attribution 4.0 International license. This license permits unrestricted use, distribution, and reproduction in any medium, provided attribution to the author(s) and the published article's title, journal citation, and DOI are maintained. Please note that some figures may have been included with permission from other third parties. It is your responsibility to obtain the proper permission from the rights holder directly for these figures.

Export citation

Export citation

Choose format for download:

Download Citation
  • Open Access

Schrödinger bridge-type diffusion models as an extension of variational autoencoders

Kentaro Kaba1,*, Reo Shimizu2, Masayuki Ohzeki1,2,3,4, and Yuki Sughiyama2

  • 1Department of Physics, Institute of Science Tokyo, Meguro-ku, Tokyo 152-9551, Japan
  • 2Graduate School of Information Sciences, Tohoku University, Sendai, Miyagi 980-9564, Japan
  • 3Research and Education Institute for Semiconductors and Informatics, Kumamoto University, Kumamoto, Kumamoto 860-8555, Japan
  • 4Sigma-i Co., Ltd., Minato-ku, Tokyo 108-0075, Japan

  • *Contact author: kaba.k.df05@m.isct.ac.jp

Phys. Rev. Research 7, 033213 – Published 3 September, 2025

DOI: https://doi.org/10.1103/dxp7-4hby

Abstract

Generative diffusion models are described by time-forward and -backward stochastic differential equations to connect the data and prior distributions. While conventional diffusion models (e.g., score-based models) only learn the backward process, more flexible frameworks have been proposed to also learn the forward process by employing the Schrödinger bridge (SB). However, due to the complexity of the mathematical structure behind SB-type models, we cannot easily give an intuitive understanding of their objective function. In this work, we propose a unified framework to construct diffusion models by reinterpreting the SB-type models as an extension of variational autoencoders. In this context, the data processing inequality plays a crucial role. As a result, we find that the objective function consists of the prior loss and drift matching parts, which enable us to reduce the numerical cost of training the forward process. Furthermore, we discuss the overfitting problem in the SB-type models in this framework.

View figure in article

Physics Subject Headings (PhySH)

Article Text

References (33)

  1. J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, Deep unsupervised learning using nonequilibrium thermodynamics, in Proceedings of the International Confer- ence on Machine Learning (PMLR, 2015), pp. 2256–2265.
  2. J. Ho, A. Jain, and P. Abbeel, Denoising diffusion probabilistic models, Adv. Neural Inf. Proc. Syst. 33, 6840 (2020).
  3. Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, Score-Based Generative Modeling Through Stochastic Differential Equations (ICLR, 2021).
  4. G. Wang, Y. Jiao, Q. Xu, Y. Wang, and C. Yang, Deep generative learning via Schrödinger bridge, in Proceedings of the 38th International Conference on Machine Learning (PMLR, 2021), Vol. 139, pp. 10794–10804.
  5. F. Vargas, P. Thodoroff, A. Lamacraft, and N. Lawrence, Solving Schrödinger bridges via maximum likelihood, Entropy 23, 1134 (2021).
  6. V. De Bortoli, J. Thornton, J. Heng, and A. Doucet, Diffusion Schrödinger bridge with applications to score-based generative modeling, Adv. Neural Inf. Proc. Syst. 34, 17695 (2021).
  7. T. Chen, G.-H. Liu, and E. A. Theodorou, Likelihood training of Schrödinger bridge using forward-backward SDEs theory, in International Conference on Learning Representations (2022).
  8. A. Tong, K. FATRAS, N. Malkin, G. Huguet, Y. Zhang, J. Rector-Brooks, G. Wolf, and Y. Bengio, Improving and generalizing flow-based generative models with minibatch optimal transport, Trans. Mach. Learn. Res. (2024).
  9. P. Dhariwal and A. Nichol, Diffusion models beat GANs on image synthesis, Adv. Neural Inf. Proc. Syst. 34, 8780 (2021).
  10. G. Batzolis, J. Stanczuk, C.-B. Schönlieb, and C. Etmann, Conditional image generation with score-based diffusion models, arXiv:2111.13606.
  11. N. Chen, Y. Zhang, H. Zen, R. J. Weiss, M. Norouzi, and W. Chan, WaveGrad: Estimating gradients for waveform generation, arXiv:2009.00713.
  12. V. Popov, I. Vovk, V. Gogoryan, T. Sadekova, and M. Kudinov, Grad-TTS: A diffusion probabilistic model for text-to-speech, in International Conference on Machine Learning (PMLR, 2021), pp. 8599–8608.
  13. Y. Song, C. Durkan, I. Murray, and S. Ermon, Maximum likelihood training of score-based diffusion models, Adv. Neural Inf. Proc. Syst. 34, 1415 (2021).
  14. H. Cao, C. Tan, Z. Gao, Y. Xu, G. Chen, P.-A. Heng, and S. Z. Li, A survey on generative diffusion models, IEEE Trans. Knowl. Data Eng. 36, 2814 (2024).
  15. E. Schrödinger, Sur la théorie relativiste de l'électron et l'interprétation de la mécanique quantique, Ann. Inst. Henri Poincare 2, 269 (1932).
  16. K. F. Caluya and A. Halder, Wasserstein proximal algorithms for the Schrödinger bridge problem: Density control with nonlinear drift, IEEE Trans. Autom. Control 67, 1163 (2021).
  17. I. Exarchos and E. A. Theodorou, Stochastic optimal control via forward and backward stochastic differential equations and importance sampling, Automatica 87, 159 (2018).
  18. P. Li, Z. Li, H. Zhang, and J. Bian, On the generalization properties of diffusion models, Adv. Neural Inf. Proc. Syst. 36, 2097 (2023).
  19. Z. Kadkhodaie, F. Guth, E. P. Simoncelli, and S. Mallat, Generalization in diffusion models arises from geometry-adaptive harmonic representation, in The Twelfth International Conference on Learning Representations (2024).
  20. D. P. Kingma and M. Welling, Auto-encoding variational Bayes, in Advances in Neural Information Processing Systems (2014).
  21. D. P. Kingma and M. Welling, An Introduction to Variational Autoencoders, Foundations and Trends® in Machine Learning 12, 307 (2019).
  22. C. Luo, Understanding diffusion models: A unified perspective, arXiv:2208.11970.
  23. T. M. Mitchell, Machine Learning (McGraw-Hill, 1997).
  24. We often choose that dimZ≤dimX and the prior is a normal distribution.
  25. This inequality usually describes the relationships among mutual informations, but in this paper, we apply it for the KL divergences.
  26. In diffusion models, dimZ=dimX.
  27. C. Gardiner, Stochastic Methods: A Handbook for the Natural and Social Sciences, Springer Series in Synergetics, 4th ed. (Springer, Berlin, 2009), Vol. 13.
  28. H. H. Risken, The Fokker-Planck Equation: Methods of Solution and Applications, Springer Series in Synergetics (Springer, 1996), 2nd ed., Vol. 18.
  29. Y. Hirono, A. Tanaka, and K. Fukushima, Understanding diffusion models by Feynman's path integral, in Proceedings of the 41st International Conference on Machine Learning (PMLR, 2024), Vol. 235, pp. 18324–18351.
  30. B. Oksendal, Stochastic Differential Equations: An Introduction with Applications (Springer Science & Business Media, 2013).
  31. A. Hyvärinen and P. Dayan, Estimation of non-normalized statistical models by score matching, J. Mach. Learn. Res. 6, 695 (2005).
  32. P. Vincent, A connection between score matching and denoising autoencoders, Neural Comput. 23, 1661 (2011).
  33. Y. Wang, L. Guo, H. Wu, and T. Zhou, Energy based diffusion generator for efficient sampling of Boltzmann distributions, arXiv:2401.02080.

Outline

Information

Sign In to Your Journals Account

Filter

Filter

Article Lookup

Enter a citation