Reuse & Permissions

It is not necessary to obtain permission to reuse this article or its components as it is available under the terms of the Creative Commons Attribution 4.0 International license. This license permits unrestricted use, distribution, and reproduction in any medium, provided attribution to the author(s) and the published article's title, journal citation, and DOI are maintained. Please note that some figures may have been included with permission from other third parties. It is your responsibility to obtain the proper permission from the rights holder directly for these figures.

Export citation

Export citation

Choose format for download:

Download Citation
  • Open Access

Reinforcement learning for real-time luminosity optimization in colliders

R. Mamutov*, G. Baranov, and A. Gerasev

  • *Contact author: R.Mamutov@inp.nsk.su

Phys. Rev. Accel. Beams 28, 122802 – Published 29 December, 2025

DOI: https://doi.org/10.1103/71jq-vbl6

Abstract

This work introduces a reinforcement learning algorithm designed for real-time luminosity optimization in collider experiments. The neural network architecture is selected from multiple candidates through systematic evaluation of their training performance. Prior to training, the input data undergo multistage preprocessing before being fed into the neural network. Our approach combines off-line pretraining on historical accelerator data with online fine-tuning during operation. By processing accelerator measurements over multisecond timescales, the reinforcement learning model dynamically adjusts the magnetic structure to maintain luminosity stability under changing beam conditions. The autonomous nature of the method eliminates the need for manual intervention, enhancing both operational efficiency and beam stability in long-term operation. Experimental validation on the VEPP-4M collider demonstrates the feasibility of the approach and provides a foundation for future development and deployment in accelerator systems.

View figure in article

Physics Subject Headings (PhySH)

Article Text

References (92)

  1. X. Huang, Beam-Based Correction and Optimization for Accelerators (Taylor & Francis, Boca Raton, 2020), 10.1201/9780429434358.
  2. R. Tomás, M. Aiba, A. Franchi, and U. Iriso, Review of linear optics measurement and correction for charged particle accelerators, Phys. Rev. Accel. Beams 20, 054801 (2017).
  3. F. Zimmermann, Measurement and correction of accelerator optics, in Beam Measurement (World Scientific, Singapore, 1999), pp. 21–107.
  4. J. Safranek, Experimental determination of storage ring optics using orbit response measurements, Nucl. Instrum. Methods Phys. Res., Sect. A 388, 27 (1997).
  5. V. Sajaev, V. Lebedev, V. Nagaslaev, and V. Valishev, Fully coupled analysis of orbit response matrices at the FNAL Tevatron, in Proceedings of the 21st Particle Accelerator Conference, Knoxville, TN, 2005 (IEEE, Piscataway, NJ, 2005), pp. 3662–3664.
  6. T. Persson, F. Carlier, J. C. De Portugal, A. G.-T. Valdivieso, A. Langner, E. H. Maclean, L. Malina, P. Skowronski, B. Salvant, R. Tomás et al., LHC optics commissioning: A journey towards 1% optics control, Phys. Rev. Accel. Beams 20, 061002 (2017).
  7. A. Romanov, D. Berkaev, P. Y. Shatunov et al., Round beam lattice correction using response matrix at VEPP-2000, in Proceedings of the International Particle Accelerator Conference, Kyoto, Japan (2010), p. 4542, https://proceedings.jacow.org/IPAC10/papers/thpe014.pdf.
  8. R. Mamutov, G. Baranov, P. Piminov, S. Sinyatkin, and D. Lipoviy, VEPP-4 M linear optics correction using orbit response matrices, J. Instrum. 19, P04026 (2024).
  9. V. Lebedev, V. Nagaslaev, A. Valishev, and V. Sajaev, Measurement and correction of linear optics and coupling at Tevatron complex, Nucl. Instrum. Methods Phys. Res., Sect. A 558, 299 (2006).
  10. H. Sugimoto, Y. Ohnishi, A. Morita, and H. Koiso, Superkekb optics tuning and issues, in Proceedings of 65th ICFA Advanced Beam Dynamics Workshop High Luminosity Circular e+e− Colliders, eeFACT’22, Frascati, Italy (2022), pp. 35–41, https://proceedings.jacow.org/eefact2022/papers/tuxat0103.pdf.
  11. A. Romanov, D. Edstrom Jr, F. Emanov, I. Koop, E. Perevedentsev, Y. A. Rogovsky, D. Shwartz, and A. Valishev, Correction of magnetic optics and beam trajectory using loco based algorithm with expanded experimental data sets, arXiv:1703.09757.
  12. Y. Alexahin and E. Gianfelice-Wendt, Determination of linear optics functions from turn-by-turn data, J. Instrum. 6, P10006 (2011).
  13. R. Jones, Measuring tune, chromaticity and coupling, arXiv:2005.02753.
  14. L. J. Wang and J. Y. Tang, Luminosity optimization and leveling in the super proton–proton collider, Radiat. Detect. Technol. Methods 5, 245 (2021).
  15. M. Dohlus, G. Hoffstaetter, M. Lomperski, and R. Wanzenberg, Report from the HERA taskforce on luminosity optimization: Theory and first luminosity scans, Technical Report No. DESY-HERA-03-01, DESY, 2003.
  16. G. Faletti, Optimisation of LHC integrated luminosity, Master’s thesis, University of Bologna, 2021.
  17. J. Ögren, A. Latina, R. Tomás, and D. Schulte, Tuning the compact linear collider 380 GeV final-focus system using realistic beam-beam signals, Phys. Rev. Accel. Beams 23, 051002 (2020).
  18. Y. Funakoshi, M. Masuzawa, K. Oide, J. Flanagan, M. Tawada, T. Ieiri, M. Tejima, M. Tobiyama, K. Ohmi, and H. Koiso, Orbit feedback system for maintaining an optimum beam collision, Phys. Rev. ST Accel. Beams 10, 101001 (2007).
  19. M. STEERING, Luminosity optimization using automated IR steering at RHIC, in Proceedings of European Particle Accelerator Conference, Lucerne, Switzerland (2004), https://proceedings.jacow.org/e04/PAPERS/MOPLT163.PDF.
  20. S. White, R. Alemany-Fernandez, M. Lamont, and H. Burkhardt, First luminosity scans in the LHC, Technical Report No. CERN-ATS-2010-096, 2010.
  21. S. White and H. Burkhardt, Luminosity optimization, in Evian 2010 Workshop on LHC Commissioning, https://cds.cern.ch/record/1281618/files/p71.pdf.
  22. S. White, Strategy for luminosity optimisation, in 2nd Evian Workshop on LHC Beam Operation (2010), https://cds.cern.ch/record/1359182/files/SW_6_03.pdf.
  23. J. Payet, A. Chancé, O. Napoly, D. Uriot, and S. Auclair, Automatic luminosity optimization of the ILC head-on beam delivery system, in Proceedings of the 2007 IEEE Particle Accelerator Conference (PAC) (IEEE, New York, 2007), pp. 1988–1990,
  24. F. Follin, R. Alemany-Fernandez, and R. Jacobsson, Online luminosity optimization at the LHC, Energy 6000, 7000 (2013), https://proceedings.jacow.org/ICALEPCS2013/papers/thppc123.pdf.
  25. M. Hostettler, R. Alemany, A. Calia, F. Follin, K. Fuchsberger, G. Hemelsoet, M. Gabriel, A. Gorzawski, M. Hruska, D. Jacquet et al., Online luminosity control and steering at the LHC, in ICALEPCS Proceedings (2017), https://cds.cern.ch/record/2305966/files/tush201.pdf.
  26. H.-F. Ji, Y. Jiao, S. Wang, D.-H. Ji, C.-H. Yu, Y. Zhang, and X.-B. Huang, Feasibility study of online tuning of the luminosity in a circular collider with the robust conjugate direction search method, Chin. Phys. C 39, 127006 (2015).
  27. X. Huang, J. Corbett, J. Safranek, and J. Wu, An algorithm for online optimization of accelerators, Nucl. Instrum. Methods Phys. Res., Sect. A 726, 77 (2013).
  28. S. Gierman, S. Ecklund, R. Field, A. Fisher, P. Grossberg, K. Krauter, E. Miller, M. Petree, K. G. Sonnad, N. Spencer et al., New fast dither system for PEP-II, in Proceedings of the European Particle Accelerator Conference (2006), p. 3029, https://proceedings.jacow.org/e06/PAPERS/THPCH100.pdf.
  29. A. S. Fisher, S. Ecklund, R. Field, S. Gierman, P. Grossberg, K. Krauter, E. Miller, M. Petree, K. Sonnad, N. Spencer et al., Commissioning the fast luminosity dither for PEP-II, in Proceedings of the 2007 IEEE Particle Accelerator Conference (PAC) (IEEE, New York, 2007), pp. 4165–4167.
  30. J. Fan and J. Wang, Online reinforcement learning control for maintaining an optimum beam collision on an electron-positron collider, Phys. Rev. Accel. Beams 27, 122801 (2024).
  31. R. Roussel, A. L. Edelen, T. Boltz, D. Kennedy, Z. Zhang, F. Ji, X. Huang, D. Ratner, A. S. Garcia, C. Xu et al., Bayesian optimization algorithms for accelerator physics, Phys. Rev. Accel. Beams 27, 084801 (2024).
  32. J. Duris, D. Kennedy, A. Hanuka, J. Shtalenkova, A. Edelen, P. Baxevanis, A. Egger, T. Cope, M. McIntire, S. Ermon et al., Bayesian optimization of a free-electron laser, Phys. Rev. Lett. 124, 124801 (2020).
  33. J. Kaiser, C. Xu, A. Eichler, A. Santamaria Garcia, O. Stein, E. Bründermann, W. Kuropka, H. Dinter, F. Mayet, T. Vinatier et al., Reinforcement learning-trained optimisers and Bayesian optimisation for online particle accelerator tuning, Sci. Rep. 14, 15733 (2024).
  34. V. Kain, S. Hirlander, B. Goddard, F. M. Velotti, G. Z. Della Porta, N. Bruchon, and G. Valentino, Sample-efficient reinforcement learning for CERN accelerator control, Phys. Rev. Accel. Beams 23, 124801 (2020).
  35. J. Kaiser, O. Stein, and A. Eichler, Learning-based optimisation of particle accelerators under partial observability without real-world training, in International Conference on Machine Learning, Baltimore, Maryland, USA (PMLR, 2022), pp. 10575–10585, https://proceedings.mlr.press/v162/kaiser22a.html.
  36. X. Chen, Y. Jia, X. Qi, Z. Wang, and Y. He, Orbit correction based on improved reinforcement learning algorithm, Phys. Rev. Accel. Beams 26, 044601 (2023).
  37. T. Boltz, M. Brosi, E. Bründermann, B. Haerer, P. Kaiser, C. Pohl, P. Schreiber, M. Yan, T. Asfour, and A.-S. Müller, Feedback design for control of the micro-bunching instability based on reinforcement learning, in CERN Yellow Reports: Conference Proceedings (2020), Vol. 9, pp. 227–227, https://proceedings.jacow.org/ipac2019/papers/mopgw017.pdf.
  38. N. Bruchon, G. Fenu, G. Gaio, M. Lonza, F. H. O’Shea, F. A. Pellegrino, and E. Salvato, Basic reinforcement learning techniques to control the intensity of a seeded free-electron laser, Electronics 9, 781 (2020).
  39. A. Awal, J. Hetzel, R. Gebel, and J. Pretz, Injection optimization at particle accelerators via reinforcement learning: From simulation to real-world application, Phys. Rev. Accel. Beams 28, 034601 (2025).
  40. D. Y. Wang, H. Bagri, C. Macdonald, S. Kiy, P. Jung, O. Shelbaya, T. Planche, W. Fedorko, R. Baartman, and O. Kester, Accelerator tuning with deep reinforcement learning, in Workshop on Machine Learning and the Physical Sciences (2021), https://ml4physicalsciences.github.io/2021/files/NeurIPS_ML4PS_2021_125.pdf.
  41. P. Piminov, G. Baranov, A. Bogomyagkov, V. Borin, V. Dorokhov, S. Karnaev, K. Y. Karyukina, V. Kiselev, E. Levichev, O. Meshkov et al., VEPP-4 M collider operation at high energy, in Proceedings of the 12th International Conference on Particle Accelerators, Campinas, SP, Brazil (2021), pp. 155–158, https://inspirehep.net/files/1a486893d0aa55d650808f2bc920437b.
  42. V. Anashin, V. Aulchenko, E. Baldin, A. Barladyan, A. Y. Barnyakov, M. Y. Barnyakov, S. Baru, I. Y. Basok, I. Bedny, O. Beloborodova et al., The KEDR detector, Phys. Part Nucl. 44, 657 (2013).
  43. E. B. Levichev, A. N. Skrinsky, Y. A. Tikhonov, and K. Y. Todyshev, High-precision particle mass measurements using the KEDR detector at the VEPP-4 M collider, Phys. Usp. 57, 66 (2014).
  44. V. Anashin, O. Anchugov, V. Aulchenko, E. Baldin, G. Baranov, A. Barladyan, A. Y. Barnyakov, M. Y. Barnyakov, S. Baru, I. Y. Basok et al., Precise measurement of Ruds and R between 1.84 and 3.72 GeV at the KEDR detector, Phys, Lett. B 788, 42 (2019).
  45. V. Anashin, V. Aulchenko, E. Baldin, A. Barladyan, A. Y. Barnyakov, M. Y. Barnyakov, S. Baru, I. Y. Basok, O. Beloborodova, A. Blinov et al., Search for narrow resonances in e+e− annihilation between 1.85 and 3.1 GeV with the KEDR detector, Phys. Lett. B 703, 543 (2011).
  46. E. Alpaydin, Introduction to Machine Learning (MIT Press, Cambridge, 2020).
  47. R. S. Sutton, A. G. Barto et al., Reinforcement Learning: An Introduction (MIT Press, Cambridge, 1998), Vol. 1.
  48. X. Pang, S. Thulasidasan, and L. Rybarcyk, Autonomous control of a particle accelerator using deep reinforcement learning, arXiv:2010.08141.
  49. S. Hirlaender and N. Bruchon, Model-free and Bayesian ensembling model-based deep reinforcement learning for particle accelerator control demonstrated on the Fermi FEL, arXiv:2012.09737.
  50. A. M. de Oñate, N. Bruchon, V. Kain, M. Schenk, and B. R. Mateos, Trajectory steering for DC beams at the CERN SPS using reinforcement learning based on intensity measurements, in Proceedings of the 16th International Particle Accelerator Conference (IPAC’25), Taipei, Taiwan (2025), https://meow.elettra.eu/81/pdf/THPM113.pdf.
  51. T. Balasooriya, S. Yoo, V. Schoefer, H.-H. Tseng, Y. Gao, W. Lin, and C. De Silva, Reinforcement learning for charged particle beam control to minimize injection mismatch in particle accelerators, in Proceedings of the 2025–2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (IEEE, New York, 2025), pp. 1–5.
  52. F. AlMahamid and K. Grolinger, Reinforcement learning algorithms: An overview and classification, in Proceedings of the 2021 IEEE Canadian Conference on Electrical and Computer Engineering (CCECE) (IEEE, New York, 2021), pp. 1–7.
  53. J. Fan, Z. Wang, Y. Xie, and Z. Yang, A theoretical analysis of Deep Q-learning, in Learning for Dynamics and Control (PMLR, 2020), pp. 486–489, https://proceedings.mlr.press/v120/yang20a.
  54. L. Graesser and W. L. Keng, Foundations of Deep Reinforcement Learning: Theory and Practice in python (Addison-Wesley Professional, 2019).
  55. V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller, Playing Atari with deep reinforcement learning, arXiv:1312.5602.
  56. M. Sewak, Deep Q network (DQN), double DQN, and dueling DQN: A step towards general artificial intelligence, in Deep Reinforcement Learning: Frontiers of Artificial Intelligence (Springer, New York, 2019), pp. 95–108.
  57. N. Sanghi, Deep Reinforcement Learning with python (Springer, New York, 2021).
  58. I. D. Mienye and T. G. Swart, A comprehensive review of deep learning: Architectures, recent advances, and applications, Information 15, 755 (2024).
  59. T. Makosso, A. Almaktoof, and K. Abo-Al-Ez, Review of different types of neural network architectures, Int. J. Electr. Eng. Appl. Sci. 7, 47 (2024).
  60. H. T. Ünal and F. Başçiftçi, Evolutionary design of neural network architectures: A review of three decades of research, Artif. Intell. Rev. 55, 1723 (2022).
  61. P. Lara-Benítez, M. Carranza-García, and J. C. Riquelme, An experimental review on deep learning architectures for time series forecasting, Int. J. Neural Syst. 31, 2130001 (2021).
  62. A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, Attention is all you need, Adv. Neural Inf. Process Syst. 30, 5998 (2017).
  63. A. Sherstinsky, Fundamentals of recurrent neural network (RNN) and long short-term memory (LSTM) network, Physica (Amsterdam) 404D, 132306 (2020).
  64. J. Kim, H. Kim, H. Kim, D. Lee, and S. Yoon, A comprehensive survey of deep learning for time series forecasting: Architectural diversity and open challenges, Artif. Intell. Rev. 58, 1 (2025).
  65. J. Sieber, C. Amo Alonso, A. Didier, M. Zeilinger, and A. Orvieto, Understanding the differences in foundation models: Attention, state space models, and recurrent neural networks, Adv. Neural Inf. Process Syst. 37, 134534 (2024). arXiv:2405.15731.
  66. F. Stolzenburg, S. Litz, O. Michael, and O. Obst, The power of linear recurrent neural networks, Neural Comput. Appl. 37, 27027 (2025).
  67. W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong et al., A survey of large language models, arXiv:2303.18223.
  68. W. Tang, G. Long, L. Liu, T. Zhou, J. Jiang, and M. Blumenstein, Rethinking 1D-CNN for time series classification: A stronger baseline, in The Tenth International Conference on Learning Representations (ICLR 2022) (2022), arXiv:2002.10061.
  69. A. Zeng, M. Chen, L. Zhang, and Q. Xu, Are transformers effective for time series forecasting?, in Proceedings of the AAAI Conference on Artificial Intelligence (2023), Vol. 37, pp. 11121–11128, https://ojs.aaai.org/index.php/AAAI/article/view/26317.
  70. Z. Ali, A comprehensive overview and comparative analysis of CNN, RNN-LSTM and transformer (2024), Available at SSRN: https://ssrn.com/abstract=5175090.
  71. E. Goman, S. Karnaev, O. Plotnikova, and E. Simonov, The database of the VEPP-4 accelerating facility parameters, in Proceedings of PCaPAC08, Ljubljana, Slovenia (JACoW, Geneva, Switzerland, 2008).
  72. The PostgreSQL Global Development Group, PostgreSQL: The world’s most advanced open source relational database, https://www.postgresql.org/.
  73. S. E. Karnaev, Control systems for the VEPP-4 accelerator complex and the NSLS-II light source booster synchrotron, Ph.D. thesis, Budker Institute of Nuclear Physics SB RAS, Novosibirsk, Russia, 2017, https://www.dissercat.com/content/sistemy-upravleniya-uskoritelnym-kompleksom-vepp-4-i-busternym-sinkhrotronom-istochnika-si.
  74. R. Mamutov and G. Baranov, Software development for automation of particle accelerator magneto-optical structure correction, Instrum. Exp. Tech. 67, S58 (2024).
  75. EPICS Community, EPICS: Experimental physics and industrial control system, https://epics-controls.org/.
  76. C. F. Worel, The curse of dimensionality: When too many features break your machine learning model, 10.2139/ssrn.5206704 (2023).
  77. I. Ramos-Pérez, J. A. Barbero-Aparicio, A. Canepa-Oneto, Á. Arnaiz-González, and J. Maudes-Raedo, An extensive performance comparison between feature reduction and feature selection preprocessing algorithms on imbalanced wide data, Information 15, 223 (2024).
  78. L. Budach, M. Feuerpfeil, N. Ihde, A. Nathansen, N. S. Noack, H. Patzlaff, H. Harmouch, and F. Naumann, The effects of data quality on ML-model performance, Inf. Syst. 132, 102549 (2025).
  79. M. A. Hall, Correlation-based feature selection for machine learning, Ph.D. thesis, The University of Waikato, 1999.
  80. M. M. Mukaka, A guide to appropriate use of correlation coefficient in medical research, Malawi Med. J. 24, 69 (2012), https://www.ajol.info/index.php/mmj/article/view/81576.
  81. F. Song, Z. Guo, and D. Mei, Feature selection using principal component analysis, in Proceedings of the 2010 International Conference on System Science, Engineering Design and Manufacturing Informatization (IEEE, New York, 2010), Vol. 1, pp. 27–30.
  82. F. Rahmat, Z. Zulkafli, A. J. Ishak, R. Z. Abdul Rahman, S. D. Stercke, W. Buytaert, W. Tahir, J. Ab Rahman, S. Ibrahim, and M. Ismail, Supervised feature selection using principal component analysis, Knowl. Inf. Syst. 66, 1955 (2024).
  83. J. Cohen, Statistical Power Analysis for the Behavioral Sciences (Routledge, New York, 2013), 10.4324/9780203771587.
  84. S. Panjeh, A. Nordahl-Hansen, and H. Cogo-Moreira, Establishing new cutoffs for Cohen’s d: An application using known effect sizes from trials for improving sleep quality on composite mental health, Int. J. Methods Psychiat. Res. 32, e1969 (2023).
  85. J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, Scaling laws for neural language models, arXiv:2001.08361.
  86. J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark et al., Training compute-optimal large language models, arXiv:2203.15556.
  87. Y. Bansal, B. Ghorbani, A. Garg, B. Zhang, C. Cherry, B. Neyshabur, and O. Firat, Data scaling laws in NMT: the effect of noise and architecture, in Proceedings of the International Conference on Machine Learning, Baltimore, Maryland, USA (PMLR, 2022), pp. 1466–1482, https://proceedings.mlr.press/v162/.
  88. N. C. Thompson, K. Greenewald, K. Lee, G. F. Manso et al., The computational limits of deep learning, arXiv:2007.05558.
  89. J. Xiong, Q. Wang, Z. Yang, P. Sun, L. Han, Y. Zheng, H. Fu, T. Zhang, J. Liu, and H. Liu, Parametrized Deep Q-networks learning: Reinforcement learning with discrete-continuous hybrid action space, arXiv:1810.06394.
  90. S. R. Sinclair, S. Banerjee, and C. L. Yu, Adaptive discretization in online reinforcement learning, Oper. Res. 71, 1636 (2023).
  91. J. Cao, L. Dong, and C. Sun, Hierarchical reinforcement learning for kinematic control tasks with parameterized action spaces, Neural Comput. Appl. 36, 323 (2024).
  92. M. Hutsebaut-Buysse, K. Mets, and S. Latré, Hierarchical reinforcement learning: A survey and open research challenges, Mach. Learn. Knowl. Extract. 4, 172 (2022).

Outline

Information

Sign In to Your Journals Account

Filter

Filter

Article Lookup

Enter a citation