Reuse & Permissions

It is not necessary to obtain permission to reuse this article or its components as it is available under the terms of the Creative Commons Attribution 4.0 International license. This license permits unrestricted use, distribution, and reproduction in any medium, provided attribution to the author(s) and the published article's title, journal citation, and DOI are maintained. Please note that some figures may have been included with permission from other third parties. It is your responsibility to obtain the proper permission from the rights holder directly for these figures.

Export citation

Export citation

Choose format for download:

Download Citation
  • Letter
  • Open Access

Deep reinforcement learning for feedback control in a collective flashing ratchet

Dong-Kyum Kim1 and Hawoong Jeong1,2,*

  • 1Department of Physics, Korea Advanced Institute of Science and Technology, Daejeon 34141, Korea
  • 2Center for Complex Systems, Korea Advanced Institute of Science and Technology, Daejeon 34141, Korea

  • *hjeong@kaist.edu

Phys. Rev. Research 3, L022002 – Published 2 April, 2021

DOI: https://doi.org/10.1103/PhysRevResearch.3.L022002

Abstract

A collective flashing ratchet transports Brownian particles using a spatially periodic, asymmetric, and time-dependent on-off switchable potential. The net current of the particles in this system can be substantially increased by feedback control based on the particle positions. Several feedback policies for maximizing the current have been proposed, but optimal policies have not been found for a moderate number of particles. Here, we use deep reinforcement learning (RL) to find optimal policies, with results showing that policies built with a suitable neural network architecture outperform the previous policies. Moreover, even in a time-delayed feedback situation where the on-off switching of the potential is delayed, we demonstrate that the policies provided by deep RL provide higher currents than the previous strategies.

View figure in article

Physics Subject Headings (PhySH)

Article Text

Supplemental Material

References (50)

  1. J. Prost, J.-F. Chauwin, L. Peliti, and A. Ajdari, Asymmetric Pumping of Particles, Phys. Rev. Lett. 72, 2652 (1994).
  2. R. D. Astumian and M. Bier, Fluctuation Driven Ratchets: Molecular Motors, Phys. Rev. Lett. 72, 1766 (1994).
  3. R. D. Astumian, Thermodynamics and kinetics of a Brownian motor, Science 276, 917 (1997).
  4. M. B. Tarlie and R. D. Astumian, Optimal modulation of a Brownian ratchet and enhanced sensitivity to a weak external force, Proc. Natl. Acad. Sci. USA 95, 2039 (1998).
  5. F. J. Cao, L. Dinis, and J. M. R. Parrondo, Feedback Control in a Collective Flashing Ratchet, Phys. Rev. Lett. 93, 040603 (2004).
  6. L. Dinis, J. M. R. Parrondo, and F. J. Cao, Closed-loop control strategy with improved current for a flashing ratchet, Europhys. Lett. 71, 536 (2005).
  7. M. Feito and F. J. Cao, Threshold feedback control for a collective flashing ratchet: Threshold dependence, Phys. Rev. E 74, 041109 (2006).
  8. M. Feito and F. J. Cao, Optimal operation of feedback flashing ratchets, J. Stat. Mech. (2009) P01031.
  9. M. Feito and F. J. Cao, Time-delayed feedback control of a flashing ratchet, Phys. Rev. E 76, 061113 (2007).
  10. E. M. Craig, B. R. Long, J. M. R. Parrondo, and H. Linke, Effect of time delay on feedback control of a flashing ratchet, Europhys. Lett. 81, 10002 (2007).
  11. E. M. Craig, N. J. Kuwada, B. J. Lopez, and H. Linke, Feedback control in flashing ratchets, Ann. Phys. 17, 115 (2008).
  12. B. J. Lopez, N. J. Kuwada, E. M. Craig, B. R. Long, and H. Linke, Realization of a Feedback Controlled Flashing Ratchet, Phys. Rev. Lett. 101, 220601 (2008).
  13. F. Roca, J. P. G. Villaluenga, and L. Dinis, Optimal protocol for a collective flashing ratchet, Europhys. Lett. 107, 10006 (2014).
  14. P. Reimann, Brownian motors: Noisy transport far from equilibrium, Phys. Rep. 361, 57 (2002).
  15. Z. Siwy and A. Fuliński, Fabrication of a Synthetic Nanopore Ion Pump, Phys. Rev. Lett. 89, 198103 (2002).
  16. I. Kosztin and K. Schulten, Fluctuation-Driven Molecular Transport Through an Asymmetric Membrane Channel, Phys. Rev. Lett. 93, 238102 (2004).
  17. O. Campàs, Y. Kafri, K. B. Zeldovich, J. Casademunt, and J.-F. Joanny, Collective Dynamics of Interacting Molecular Motors, Phys. Rev. Lett. 97, 038101 (2006).
  18. J. Brugués and J. Casademunt, Self-Organization and Cooperativity of Weakly Coupled Molecular Motors under Unequal Loading, Phys. Rev. Lett. 102, 118104 (2009).
  19. D. Oriola and J. Casademunt, Cooperative Force Generation of KIF1A Brownian Motors, Phys. Rev. Lett. 111, 048103 (2013).
  20. W. Hwang and M. Karplus, Structural basis for power Stroke vs. Brownian ratchet mechanisms of motor proteins, Proc. Natl. Acad. Sci. USA 116, 19777 (2019).
  21. I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning (MIT Press, Cambridge, MA, 2016).
  22. V. Bapst, T. Keck, A. Grabska-Barwińska, C. Donner, E. D. Cubuk, S. S. Schoenholz, A. Obika, A. W. R. Nelson, T. Back, D. Hassabis, and P. Kohli, Unveiling the predictive power of static structure in glassy systems, Nat. Phys. 16, 448 (2020).
  23. J. Carrasquilla, Machine learning for quantum matter, Adv. Phys.: X 5, 1797528 (2020).
  24. G. Carleo, I. Cirac, K. Cranmer, L. Daudet, M. Schuld, N. Tishby, L. Vogt-Maranto, and L. Zdeborová, Machine learning and the physical sciences, Rev. Mod. Phys. 91, 045002 (2019).
  25. R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction (MIT Press, Cambridge, MA, 2018).
  26. V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al., Human-level control through deep reinforcement learning, Nature (London) 518, 529 (2015).
  27. D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot et al., Mastering the game of Go with deep neural networks and tree search, Nature (London) 529, 484 (2016).
  28. D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, T. Lillicrap, K. Simonyan, and D. Hassabis, A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play, Science 362, 1140 (2018).
  29. O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev et al., Grandmaster level in StarCraft II using multi-agent reinforcement learning, Nature (London) 575, 350 (2019).
  30. T. Fösel, P. Tighineanu, T. Weiss, and F. Marquardt, Reinforcement Learning with Neural Networks for Quantum Feedback, Phys. Rev. X 8, 031084 (2018).
  31. R. Porotti, D. Tamascelli, M. Restelli, and E. Prati, Coherent transport of quantum states by deep reinforcement learning, Commun. Phys. 2, 61 (2019).
  32. M. Y. Niu, S. Boixo, V. N. Smelyanskiy, and H. Neven, Universal quantum control through deep reinforcement learning, npj Quantum Inf. 5, 33 (2019).
  33. Z. An and D. L. Zhou, Deep reinforcement learning for quantum gate control, Europhys. Lett. 126, 60002 (2019).
  34. Z. T. Wang, Y. Ashida, and M. Ueda, Deep Reinforcement Learning Control of Quantum Cartpoles, Phys. Rev. Lett. 125, 100401 (2020).
  35. J. Achiam, Spinning Up in Deep Reinforcement Learning, 2018, https://spinningup.openai.com.
  36. J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, Proximal policy optimization algorithms, arXiv:1707.06347.
  37. See Supplemental Material at http://link.aps.org/supplemental/10.1103/PhysRevResearch.3.L022002 for the training details, hyperparameters, neural network architecture configurations, policy and value networks over time, the results on the sawtooth potential, and the source code for the runs and results, which includes Refs. [35, 36, 47, 48, 49, 50].
  38. M. Zaheer, S. Kottur, S. Ravanbakhsh, B. Poczos, R. R. Salakhutdinov, and A. J. Smola, Deep sets, in Advances in Neural Information Processing Systems 30 (Curran Associates, Long Beach, CA, 2017), pp. 3391–3401.
  39. K. V. Katsikopoulos and S. E. Engelbrecht, Markov decision processes with delays and asynchronous cost collection, IEEE Trans. Autom. Control 48, 568 (2003).
  40. K. Cho, B. van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, Learning phrase representations using RNN encoder–decoder for statistical machine translation, in Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) (Association for Computational Linguistics, Doha, Qatar, 2014), pp. 1724–1734.
  41. M. Feito, J. P. Baltanás, and F. J. Cao, Rocking feedback-controlled ratchets, Phys. Rev. E 80, 031128 (2009).
  42. F. J. Cao, M. Feito, and H. Touchette, Information and flux in a feedback controlled Brownian ratchet, Physica A 388, 113 (2009).
  43. F. J. Cao and M. Feito, Thermodynamics of feedback controlled systems, Phys. Rev. E 79, 041118 (2009).
  44. T. Sagawa and M. Ueda, Nonequilibrium thermodynamics of feedback control, Phys. Rev. E 85, 021104 (2012).
  45. J. M. R. Parrondo, J. M. Horowitz, and T. Sagawa, Thermodynamics of information, Nat. Phys. 11, 131 (2015).
  46. G. Dulac-Arnold, N. Levine, D. J. Mankowitz, J. Li, C. Paduraru, S. Gowal, and T. Hester, An empirical investigation of the challenges of real-world reinforcement learning, arXiv:2003.11881.
  47. A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al., PyTorch: An imperative style, high-performance deep learning library, in Advances in Neural Information Processing Systems 32 (Curran Associates, Vancouver, 2019), pp. 8024–8035.
  48. V. Nair and G. E. Hinton, Rectified linear units improve restricted Boltzmann machines, in Proceedings of the 27th International Conference on Machine Learning, Haifa, Israel (Omnipress, Madison, WI, 2010), pp. 807–814.
  49. D. P. Kingma and J. Ba, Adam: A method for stochastic optimization, in International Conference on Learning Representations (2015), arXiv:1412.6980.
  50. J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel, High-dimensional continuous control using generalized advantage estimation, in International Conference on Learning Representations (2016), arXiv:1506.02438.

Outline

Information

Sign In to Your Journals Account

Filter

Filter

Article Lookup

Enter a citation