- Open Access
General-purpose quantum architecture search based on deep reinforcement learning
Phys. Rev. A 112, 052409 – Published 4 November, 2025
DOI: https://doi.org/10.1103/7rc4-p446
Abstract
Reinforcement learning (RL) shows promise for automated quantum circuit design but often stalls due to a “fidelity trap”: by optimizing only state fidelity, agents overlook entanglement structure and become stranded in suboptimal, overly complex circuits, resulting in significantly reduced search efficiency. In this paper, we propose a scheme that overcomes this barrier by implementing an entanglement-aware learning framework and enhancing the agent's reward function with a direct, quantitative measure of entanglement. This approach offers a more comprehensive physical description of the state space. We demonstrate the efficacy of this principle on three- and four-qubit state-synthesis tasks within an expanded gate set. For this problem, where the fidelity-driven agent systematically fails to discover the minimal-depth circuit, our entanglement-aware agent consistently succeeds. This transformative result is highly robust against variations in initial random seeds and extends to multiqubit systems even in the presence of noise. Our findings establish a generalizable principle that incorporating entanglement as an auxiliary reward can significantly enhance RL-based solutions for a broad class of fidelity-centric tasks in quantum physics and pave the way for scalable, automated discovery on near-term quantum devices.
Physics Subject Headings (PhySH)
Article Text
References (62)
- H.-K. Lau and M. B. Plenio, Universal quantum computing with arbitrary continuous-variable encoding, Phys. Rev. Lett. 117, 100501 (2016).
- I. M. Georgescu, S. Ashhab, and F. Nori, Quantum simulation, Rev. Mod. Phys. 86, 153 (2014).
- M. D'Arcangelo, L.-P. Henry, L. Henriet, D. Loco, N. Gouraud, S. Angebault, J. Sueiro, J. Forêt, P. Monmarché, and J.-P. Piquemal, Leveraging analog quantum computing with neutral atoms for solvent configuration prediction in drug discovery, Phys. Rev. Res. 6, 043020 (2024).
- R. Babbush, N. Wiebe, J. McClean, J. McClain, H. Neven, and G. K.-L. Chan, Low-depth quantum simulation of materials, Phys. Rev. X 8, 011044 (2018).
- M. Cerezo, A. Arrasmith, R. Babbush, S. C. Benjamin, S. Endo, K. Fujii, J. R. McClean, K. Mitarai, X. Yuan, L. Cincio, and P. J. Coles, Variational quantum algorithms, Nat. Rev. Phys. 3, 625 (2021).
- A. Kundu, P. Bedełek, M. Ostaszewski, O. Danaci, Y. J. Patel, V. Dunjko, and J. A. Miszczak, Enhancing variational quantum state diagonalization using reinforcement learning techniques, New J. Phys. 26, 013034 (2024).
- E. Ye and S. Y.-C. Chen, Quantum architecture search via continual reinforcement learning, arXiv:2112.05779.
- Y. Du, T. Huang, S. You, M.-H. Hsieh, and D. Tao, Quantum circuit architecture search for variational quantum algorithms, npj Quantum Inf. 8, 62 (2022).
- W. Zhu, J. Pi, and Q. Peng, A brief survey of quantum architecture search, in Proceedings of the 6th International Conference on Algorithms, Computing and Systems, ICACS'22 (Association for Computing Machinery, New York, 2023).
- D. Martyniuk, J. Jung, and A. Paschke, Quantum architecture search: A survey, in 2024 IEEE International Conference on Quantum Computing and Engineering (QCE), Vol. 1 (IEEE, New York, 2024), pp. 1695–1706.
- Y. J. Patel, A. Kundu, M. Ostaszewski, X. Bonet-Monroig, V. Dunjko, and O. Danaci, Curriculum reinforcement learning for quantum architecture search under hardware errors, arXiv:2402.03500.
- S. Y.-C. Chen, Quantum reinforcement learning for quantum architecture search, in Proceedings of the 2023 International Workshop on Quantum Classical Cooperative, QCCC'23 (Association for Computing Machinery, New York, 2023), pp. 17–20.
- S.-X. Zhang, C.-Y. Hsieh, S. Zhang, and H. Yao, Differentiable quantum architecture search, Quantum Sci. Technol. 7, 045023 (2022).
- X. Zhu and X. Hou, Quantum architecture search via truly proximal policy optimization, Sci. Rep. 13, 5157 (2023).
- Y. Sun, Z. Wu, Y. Ma, and V. Tresp, Quantum architecture search with unsupervised representation learning, arXiv:2401.11576.
- E.-J. Kuo, Y.-L. L. Fang, and S. Y.-C. Chen, Quantum architecture search via deep reinforcement learning, arXiv:2104.07715.
- R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction (MIT, Cambridge, 2018).
- J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, Proximal policy optimization algorithms, arXiv:1707.06347.
- V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Harley, T. P. Lillicrap, D. Silver, and K. Kavukcuoglu, Asynchronous methods for deep reinforcement learning, in Proceedings of the 33rd International Conference on International Conference on Machine Learning, edited by M. F. Balcan and K. Q. Weinberger, Proceedings of Machine Learning Research Vol. 48 (PMLR, New York, NY, 2016), pp. 1928–1937.
- V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller, Playing Atari with deep reinforcement learning, arXiv:1312.5602.
- J. Fan, A review for deep reinforcement learning in Atari: Benchmarks, challenges, and solutions, arXiv:2112.04145.
- L. Kaiser, M. Babaeizadeh, P. Miłos, B. Osiński, R. H. Campbell, K. Czechowski, D. Erhan, C. Finn, P. Kozakowski, S. Levine, A. Mohiuddin, R. Sepassi, G. Tucker, and H. Michalewski, Model-based reinforcement learning for Atari, in 8th International Conference on Learning Representations, Addis Ababa, Ethiopia (OpenReview.net, 2020).
- V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al., Human-level control through deep reinforcement learning, Nature (London) 518, 529 (2015).
- H. Van Hasselt, A. Guez, and D. Silver, Deep reinforcement learning with double Q-learning, in Proceedings of the AAAI Conference on Artificial Intelligence (AAAI Press, Palo Alto, CA, 2016).
- W. Ye, S. Liu, T. Kurutach, P. Abbeel, and Y. Gao, Mastering Atari games with limited data, 35th Conference on Neural Information Processing Systems (NeurIPS) (Curran Associates, New York, 2021), Vol. 34, pp. 25476–25488.
- B. M. Babayan, N. Uchida, and S. J. Gershman, Belief state representation in the dopamine system, Nat. Commun. 9, 1891 (2018).
- K. Arulkumaran, M. P. Deisenroth, M. Brundage, and A. A. Bharath, Deep reinforcement learning: A brief survey, IEEE Signal Proces. Mag. 34, 26 (2017).
- P. Andreasson, J. Johansson, S. Liljestrand, and M. Granath, Quantum error correction for the toric code using deep reinforcement learning, Quantum 3, 183 (2019).
- D. Fitzek, M. Eliasson, A. F. Kockum, and M. Granath, Deep Q-learning decoder for depolarizing noise on the toric code, Phys. Rev. Res. 2, 023230 (2020).
- H. P. Nautrup, N. Delfosse, V. Dunjko, H. J. Briegel, and N. Friis, Optimizing quantum error correction codes with reinforcement learning, Quantum 3, 215 (2019).
- Y. Zeng, Z.-Y. Zhou, E. Rinaldi, C. Gneiting, and F. Nori, Approximate autonomous quantum error correction with reinforcement learning, Phys. Rev. Lett. 131, 050601 (2023).
- L. Domingo Colomer, M. Skotiniotis, and R. Muñoz-Tapia, Reinforcement learning for optimal error correction of toric codes, Phys. Lett. A 384, 126353 (2020).
- S. Z. Baba, N. Yoshioka, Y. Ashida, and T. Sagawa, Deep reinforcement learning for preparation of thermal and prethermal quantum states, Phys. Rev. Appl. 19, 014068 (2023).
- X. Zhao, Y. Zhao, M. Li, T. Li, Q. Liu, S. Guo, and X. Yi, A strategy for preparing quantum squeezed states using reinforcement learning, Ann. Phys. 536, 2400056 (2024).
- R. Porotti, A. Essig, B. Huard, and F. Marquardt, Deep reinforcement learning for quantum state preparation with weak nonlinear measurements, Quantum 6, 747 (2022).
- R. H. He, R. Wang, S. S. Nie, and Z. M. Wang, Deep reinforcement learning for universal quantum state preparation via dynamic pulse control, EPJ Quantum Technol. 8, 29 (2021).
- X.-M. Zhang, Z. Wei, R. Asad, X.-C. Yang, and X. Wang, When does reinforcement learning stand out in quantum control? A comparative study on state preparation, npj Quantum Inf. 5, 85 (2019).
- S. Borah, B. Sarma, M. Kewming, G. J. Milburn, and J. Twamley, Measurement-based feedback quantum control with deep reinforcement learning for a double-well nonlinear potential, Phys. Rev. Lett. 127, 190403 (2021).
- V. V. Sivak, A. Eickbusch, H. Liu, B. Royer, I. Tsioutsios, and M. H. Devoret, Model-free quantum control with reinforcement learning, Phys. Rev. X 12, 011059 (2022).
- M. Bukov, A. G. R. Day, D. Sels, P. Weinberg, A. Polkovnikov, and P. Mehta, Reinforcement learning in different phases of quantum control, Phys. Rev. X 8, 031086 (2018).
- Y. Zeng, J. Shen, S. Hou, T. Gebremariam, and C. Li, Quantum control based on machine learning in an open quantum system, Phys. Lett. A 384, 126886 (2020).
- T. Neema, S. Jha, and T. Sahai, Non-Markovian quantum control via model maximum likelihood estimation and reinforcement learning, arXiv:2402.05084.
- Z. An and D. L. Zhou, Deep reinforcement learning for quantum gate control, Europhys. Lett. 126, 60002 (2019).
- Y. Baum, M. Amico, S. Howell, M. Hush, M. Liuzzi, P. Mundada, T. Merkh, A. R. R. Carvalho, and M. J. Biercuk, Experimental deep reinforcement learning for error-robust gate-set design on a superconducting quantum computer, PRX Quantum 2, 040324 (2021).
- H. Nam Nguyen, F. Motzoi, M. Metcalf, K. Birgitta Whaley, M. Bukov, and M. Schmitt, Reinforcement learning pulses for transmon qubit entangling gates, Mach. Learn.: Sci. Technol. 5, 025066 (2024).
- R. J. Williams, Simple statistical gradient-following algorithms for connectionist reinforcement learning, Mach. Learn. 8, 229 (1992).
- F. Pardo, A. Tavakoli, V. Levdik, and P. Kormushev, Time limits in reinforcement learning, in Proceedings of the 35th International Conference on Machine Learning, Vol. 80, edited by J. Dy and A. Krause (PMLR, Cambridge, MA, 2018), pp. 4045–4054.
- D. Arumugam, P. Henderson, and P.-L. Bacon, An information-theoretic perspective on credit assignment in reinforcement learning, arXiv:2103.06224.
- S. Singh, R. L. Lewis, and A. G. Barto, Where do rewards come from, Proc. Ann. Meet. Cogn. Sci. Soc. 49, 2601 (2009).
- D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané, Concrete problems in AI safety, arXiv:1606.06565.
- D. Dewey, Reinforcement learning and the reward engineering principle, in AAAI Spring Symposia (AAAI Press, Palo Alto, CA, 2014).
- J. G. G. de Oliveira, Jr., J. G. Peixoto de Faria, and M. C. Nemes, Residual entanglement and sudden death: A direct connection, Phys. Lett. A 375, 4255 (2011).
- K. Cao, Z.-W. Zhou, G.-C. Guo, and L. He, Efficient numerical method to calculate the three-tangle of mixed states, Phys. Rev. A 81, 034302 (2010).
- V. Coffman, J. Kundu, and W. K. Wootters, Distributed entanglement, Phys. Rev. A 61, 052306 (2000).
- M. Ma, Y. Li, and J. Shang, Multipartite entanglement measures: A review, Fundam. Res. 55, 145303 (2024).
- J. de Jong, F. Hahn, N. Tcholtchev, M. Hauswirth, and A. Pappa, Extracting GHZ states from linear cluster states, Phys. Rev. Res. 6, 013330 (2024).
- R. Islam, R. Ma, P. M. Preiss, M. Eric Tai, A. Lukin, M. Rispoli, and M. Greiner, Measuring entanglement entropy in a quantum many-body system, Nature (London) 528, 77 (2015).
- Y.-H. Liu, Q.-S. Tan, L.-M. Kuang, and J.-Q. Liao, Deterministic generation of nonclassical mechanical states in cavity optomechanics via reinforcement learning, Phys. Rev. A 111, 053517 (2025).
- H. Wang, Y. Ding, J. Gu, Y. Lin, D. Z. Pan, F. T. Chong, and S. Han, Quantumnas: Noise-adaptive search for robust quantum circuits, in 2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA) (IEEE, New York, 2022), pp. 692–708.
- Y. Huang, Q. Li, X. Hou, R. Wu, M.-H. Yung, A. Bayat, and X. Wang, Robust resource-efficient quantum variational ansatz through an evolutionary algorithm, Phys. Rev. A 105, 052414 (2022).
- Z. Xu, W. Shang, S. Kim, A. Bobbitt, E. Lee, and T. Luo, Quantum-inspired genetic algorithm for designing planar multilayer photonic structure, NPJ Comput. Mater. 10, 257 (2024).
- X. Bi, Code for “General-purpose quantum architecture search based on deep reinforcement learning (version 1.0.0) (2025)”, https://github.com/vojhs/QAS-entanglement.