- Open Access
Accelerated Deep Atomic Potential Transformer for long-range machine learning force fields without graphs
Phys. Rev. Research 8, 033266 – Published 3 September, 2026
DOI: https://doi.org/10.1103/3xf3-ts5f
Abstract
Point defects play a central role in driving the properties of materials. First-principles methods are widely used to compute defect energetics and structures, including at scale for high-throughput defect databases. However, these methods are computationally expensive, making machine learning force fields (MLFFs) an attractive alternative for accelerating structural relaxations. Most existing MLFFs are based on graph neural networks (GNNs), which can suffer from oversmoothing, oversquashing, and poor representation of long-range interactions. Both of these issues are especially of concern when modeling point defects. To address these challenges, we introduce the Accelerated Deep Atomic Potential Transformer (ADAPT), an MLFF that replaces graph representations with a direct coordinates-in-space formulation and explicitly considers all pairwise atomic interactions. Atoms are treated as “tokens,” with a transformer encoder modeling their interactions. Applied to a dataset of silicon point defects, ADAPT achieves an reduction in force and an reduction in energy prediction error relative to a state-of-the-art GNN-based model, while requiring only a fraction of the computational cost.
Physics Subject Headings (PhySH)
Article Text
References (98)
- M. M. Bronstein, J. Bruna, T. Cohen, and P. Veličković, Geometric deep learning: Grids, groups, graphs, geodesics, and gauges, arXiv:2104.13478.
- P. Reiser, M. Neubert, A. Eberhard, L. Torresi, C. Zhou, C. Shao, H. Metni, C. van Hoesel, H. Schopmans, T. Sommer, et al., Graph neural networks for materials science and chemistry, Commun. Mater. 3, 93 (2022).
- I. Batatia, D. P. Kovacs, G. Simm, C. Ortner, and G. Csányi, MACE: Higher order equivariant message passing neural networks for fast and accurate force fields, Adv. Neural Inf. Process. Syst. 35, 11423 (2022).
- B. Deng, P. Zhong, K. Jun, J. Riebesell, K. Han, C. J. Bartel, and G. Ceder, CHGNet as a pretrained universal neural network potential for charge-informed atomistic modelling, Nat. Mach. Intell. 5, 1031 (2023).
- H. Yang, C. Hu, Y. Zhou, X. Liu, Y. Shi, J. Li, G. Li, Z. Chen, S. Chen, C. Zeni, et al., MatterSim: A deep learning atomistic model across elements, temperatures and pressures, arXiv:2405.04967.
- J. T. Frank, O. T. Unke, K.-R. Müller, and S. Chmiela, A Euclidean transformer for fast and stable machine learned force fields, Nat. Commun. 15, 6539 (2024).
- I. Poltavsky and A. Tkatchenko, Machine learning force fields: Recent advances and remaining challenges, J. Phys. Chem. Lett. 12, 6551 (2021).
- C. Chen and S. P. Ong, A universal graph deep learning interatomic potential for the periodic table, Nat. Comput. Sci. 2, 718 (2022).
- K. Choudhary and B. DeCost, Atomistic line graph neural network for improved materials property predictions, npj Comput. Mater. 7, 185 (2021).
- K. Schütt, P.-J. Kindermans, H. E. Sauceda Felix, S. Chmiela, A. Tkatchenko, and K.-R. Müller, SchNet: A continuous-filter convolutional neural network for modeling quantum interactions, Adv. Neural Inf. Process. Syst. 30 (2017).
- S. Batzner, A. Musaelian, L. Sun, M. Geiger, J. P. Mailoa, M. Kornbluth, N. Molinari, T. E. Smidt, and B. Kozinsky, E(3)-equivariant graph neural networks for data-efficient and accurate interatomic potentials, Nat. Commun. 13, 2453 (2022).
- A. Musaelian, S. Batzner, A. Johansson, L. Sun, C. J. Owen, M. Kornbluth, and B. Kozinsky, Learning local equivariant representations for large-scale atomistic dynamics, Nat. Commun. 14, 579 (2023).
- S. T. Miller, J. F. Lindner, A. Choudhary, S. Sinha, and W. L. Ditto, The scaling of physics-informed machine learning with data and dimensions, Chaos Solitons Fractals: X 5, 100046 (2020).
- F. Hendriks, V. Menkovski, M. Doškář, M. G. Geers, and O. Rokoš, Similarity equivariant graph neural networks for homogenization of metamaterials, Comput. Methods Appl. Mech. Eng. 439, 117867 (2025).
- C. Vignac, A. Loukas, and P. Frossard, Building powerful and equivariant graph neural networks with structural message-passing, Adv. Neural Inf. Process. Syst. 33, 14143 (2020).
- J. T. Frank, O. T. Unke, and K.-R. Müller, So3krates: Equivariant attention for interactions on arbitrary length-scales in molecular systems, Adv. Neural Inf. Proc. Systems 35, 29400 (2022).
- M. H. Rahman, P. Gollapalli, P. Manganaris, S. K. Yadav, G. Pilania, B. DeCost, K. Choudhary, and A. Mannodi-Kanakkithodi, Accelerating defect predictions in semiconductors using graph neural networks, APL Mach. Learn. 2, 016122 (2024).
- X. Xiang, D. Soh, and S. Dunham, Exploration of deep learning models for accelerated defect property predictions and device design of cubic semiconductor crystals, J. Phys. Chem. C 128, 8821 (2024).
- Z. Fang and Q. Yan, Leveraging persistent homology features for accurate defect formation energy predictions via graph neural networks, Chem. Mater. 37, 1531 (2025).
- I. Mosquera-Lois, S. R. Kavanagh, A. M. Ganose, and A. Walsh, Machine-learning structural reconstructions for accelerated point defect calculations, npj Comput. Mater. 10, 121 (2024).
- C. Freysoldt, B. Grabowski, T. Hickel, J. Neugebauer, G. Kresse, A. Janotti, and C. G. Van de Walle, First-principles calculations for point defects in solids, Rev. Mod. Phys. 86, 253 (2014).
- V. Ivády, I. A. Abrikosov, and A. Gali, First principles calculation of spin-related quantities for point defect qubit research, npj Comput. Mater. 4, 76 (2018).
- Q. Li, Z. Han, and X.-M. Wu, Deeper insights into graph convolutional networks for semi-supervised learning, in Proceedings of the AAAI Conference on Artificial Intelligence, (2018), Vol. 32.
- Z. Yang, X. Liu, X. Zhang, P. Huang, K. S. Novoselov, and L. Shen, Modeling crystal defects using defect informed neural networks, npj Comput. Mater. 11, 229 (2025).
- A. D. Lopez-Rojas and C. A. Cruz-Villar, Neural networks as an approximator for a family of optimization algorithm solutions for online applications, Neural Comput. Appl. 36, 3125 (2024).
- B. Amos, Tutorial on amortized optimization, Found. Trends Mach. Learn. 16, 592 (2023).
- R. Qiu, Z. Sun, and Y. Yang, DIMES: A differentiable meta solver for combinatorial optimization problems, Adv. Neural Inf. Process. Syst. 35, 25531 (2022).
- A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, Attention is all you need, Adv. Neural Inf. Process. Syst. 30 (2017).
- C. Zhou, Q. Li, C. Li, J. Yu, Y. Liu, G. Wang, K. Zhang, C. Ji, Q. Yan, L. He, et al., A comprehensive survey on pretrained foundation models: A history from BERT to ChatGPT, Int. J. Mach. Learn. Cybern. 1 (2024).
- A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al., An image is worth 16 × 16 words: Transformers for image recognition at scale, arXiv:2010.11929.
- J. Abramson, J. Adler, J. Dunger, R. Evans, T. Green, A. Pritzel, O. Ronneberger, L. Willmore, A. J. Ballard, J. Bambrick, et al., Accurate structure prediction of biomolecular interactions with AlphaFold 3, Nature (London) 630, 493 (2024).
- Y. Xiong, C. Bourgois, N. Sheremetyeva, W. Chen, D. Dahliah, H. Song, J. Zheng, S. M. Griffin, A. Sipahigil, and G. Hautier, High-throughput identification of spin-photon interfaces in silicon, Sci. Adv. 9, eadh8617 (2023).
- Y. Xiong, J. Zheng, S. McBride, X. Zhang, S. M. Griffin, and G. Hautier, Computationally driven discovery of T center-like quantum defects in silicon, J. Am. Chem. Soc. 146, 30046 (2024).
- T. Kreiman, Y. Bai, F. Atieh, E. Weaver, E. Qu, and A. S. Krishnapriyan, Transformers discover molecular structure without graph priors, arXiv:2510.02259.
- A. A. Elhag, A. Raja, A. Morehead, S. M. Blau, G. M. Morris, and M. M. Bronstein, Learning inter-atomic potentials without explicit equivariance, arXiv:2510.00027.
- M. Eissler, T. Korjakow, S. Ganscha, O. T. Unke, K.-R. Müller, and S. Gugler, How simple can you go? An off-the-shelf transformer approach to molecular dynamics, J. Chem. Phys. 164, 094308 (2026).
- S. Chmiela, A. Tkatchenko, H. E. Sauceda, I. Poltavsky, K. T. Schütt, and K.-R. Müller, Machine learning of accurate energy-conserving molecular force fields, Sci. Adv. 3, e1603015 (2017).
- D. Zhang, H. Bi, F.-Z. Dai, W. Jiang, X. Liu, L. Zhang, and H. Wang, Pretraining of attention-based deep learning potential model for molecular simulation, npj Comput. Mater. 10, 94 (2024).
- F. Bigi, M. Langer, and M. Ceriotti, The dark side of the forces: Assessing non-conservative force models for atomistic machine learning, in Proceedings of the 42nd International Conference on Machine Learning, edited by A. Singh, M. Fazel, D. Hsu, S. Lacoste-Julien, F. Berkenkamp, T. Maharaj, K. Wagstaff, and J. Zhu, Proceedings of Machine Learning Research, Vol. 267 (PMLR, 2025), pp. 4384–4414.
- X. Fu, Z. Wu, W. Wang, T. Xie, S. Keten, R. Gomez-Bombarelli, and T. Jaakkola, Forces are not enough: Benchmark and critical evaluation for machine learning force fields with molecular simulations, arXiv:2210.07237.
- M. Neumann, J. Gin, B. Rhodes, S. Bennett, Z. Li, H. Choubisa, A. Hussey, and J. Godwin, Orb: A fast, scalable neural network potential, arXiv:2410.22570.
- B. Rhodes, S. Vandenhaute, V. Šimkus, J. Gin, J. Godwin, T. Duignan, and M. Neumann, Orb-v3: Atomistic simulation at scale, arXiv:2504.06231.
- K. T. Schütt, H. E. Sauceda, P.-J. Kindermans, A. Tkatchenko, and K.-R. Müller, SchNet—A deep learning architecture for molecules and materials, J. Chem. Phys. 148, 241722 (2018).
- O. T. Unke, S. Chmiela, H. E. Sauceda, M. Gastegger, I. Poltavsky, K. T. Schutt, A. Tkatchenko, and K.-R. Müller, Machine learning force fields, Chem. Rev. 121, 10142 (2021).
- J. J. Webster and C. Kit, Tokenization as the initial phase in NLP, in Proceedings of the 14th conference on Computational linguistics (COLING '92) (Association for Computational Linguistics, Nantes, France, 1992), Vol. 4, pp. 1106–1110.
- J. P. Perdew, K. Burke, and M. Ernzerhof, Generalized gradient approximation made simple, Phys. Rev. Lett. 77, 3865 (1996).
- G. Cybenko, Approximation by superpositions of a sigmoidal function, Math. Control Signal Systems 2, 303 (1989).
- N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, Dropout: A simple way to prevent neural networks from overfitting, J. Mach. Learn. Res. 15, 1929 (2014).
- An alternative weighting would be (F1)We experimented with Eq. (F1), but found that for silicon defects Eq. (12) gave better results. It is possible that Eq. (F1) would perform better in some applications.
- D. Jha, L. Ward, A. Paul, W.-k. Liao, A. Choudhary, C. Wolverton, and A. Agrawal, ElemNet: Deep learning the chemistry of materials from only elemental composition, Sci. Rep. 8, 17593 (2018).
- Y. Liang, M. Chen, Y. Wang, H. Jia, T. Lu, F. Xie, G. Cai, Z. Wang, S. Meng, and M. Liu, A universal model for accurately predicting the formation energy of inorganic compounds, Sci. China Mater. 66, 343 (2023).
- L. Zhang, J. Han, H. Wang, R. Car, and W. E, Deep potential molecular dynamics: A scalable model with the accuracy of quantum mechanics, Phys. Rev. Lett. 120, 143001 (2018).
- J. Müller, On the space-time expressivity of ResNets, arXiv:1910.09599.
- J. Baggenstos and D. Salimova, Approximation properties of residual neural networks for Kolmogorov PDEs, arXiv:2111.00215.
- M. M. Moghaddam, K. Parand, and S. R. Kheradpisheh, Advanced physics-informed neural network with residuals for solving complex integral equations, Anal. Num. Solutions Nonlinear Equat. 11, 153 (2026).
- A. Noorizadegan, R. Cavoretto, D.-L. Young, and C.-S. Chen, Stable weight updating: A key to reliable PDE solutions using deep learning, Eng. Anal. Boundary Elem. 168, 105933 (2024).
- K. Kashinath, M. Mustafa, A. Albert, J. Wu, C. Jiang, S. Esmaeilzadeh, K. Azizzadenesheli, R. Wang, A. Chattopadhyay, A. Singh et al., Physics-informed machine learning: Case studies for weather and climate modelling, Philos. Trans. R. Soc. A 379, 20200093 (2021).
- I. Batatia, P. Benner, Y. Chiang, A. M. Elena, D. P. Kovács, J. Riebesell, X. R. Advincula, M. Asta, M. Avaylon, W. J. Baldwin et al., A foundation model for atomistic materials chemistry, J. Chem. Phys. 163, 184110 (2025).
- S. Maheshwari, Z. Tang, J. Ock, A. Kolluru, A. B. Farimani, and J. R. Kitchin, Beyond force metrics: Pre-training MLFFs for stable MD simulations, arXiv:2506.14850.
- E. Bitzek, P. Koskinen, F. Gähler, M. Moseler, and P. Gumbsch, Structural relaxation made simple, Phys. Rev. Lett. 97, 170201 (2006).
- The relaxation uses a version of MACE, which achieved a MAE of 0.0217, rather than the most recent retraining with error 0.0161. However, the difference in force accuracy is small compared to the difference in relaxation quality.
- A. Alkauskas, B. B. Buckley, D. D. Awschalom, and C. G. Van de Walle, First-principles theory of the luminescence lineshape for the triplet transition in diamond NV centres, New J. Phys. 16, 073026 (2014).
- E. Dramko, Y. Zhu, A. Krivokapic, G. Hautier, T. Reps, C. Jermaine, and A. Kyrillidis, On the finetuning of MLIPs through the lens of iterated maps with BPTT, arXiv:2512.01067.
- The authors successfully trained ADAPT with 216 atom supercells on a 2023 M2 MacBook Air laptop.
- S. Liang, Y. Wang, C. Liu, L. He, H. Li, D. Xu, and X. Li, EnGN: A high-throughput and energy-efficient accelerator for large graph neural networks, IEEE Trans. Comput. 70, 1511 (2020).
- R. Sutton, The bitter lesson, Incomplete Ideas (blog) 13, 38 (2019).
- W. Deng and J. Rao, MEGA: More efficient graph attention for GNNs, in 2024 IEEE 44th International Conference on Distributed Computing Systems (ICDCS), Jersey City, NJ, USA (IEEE, 2024), pp. 71–81.
- A. Auten, M. Tomei, and R. Kumar, Hardware acceleration of graph neural networks, in 2020 57th ACM/IEEE Design Automation Conference (DAC), San Francisco, CA, USA (IEEE, 2020), pp. 1–6.
- Y. Zhang and Q. Yang, A survey on multi-task learning, IEEE Trans. Knowl. Data Eng. 34, 5586 (2021).
- Y. Liu, E. Sangineto, W. Bi, N. Sebe, B. Lepri, and M. Nadai, Efficient training of visual transformers with small datasets, Adv. Neural Inf. Process. Syst. 34, 23818 (2021).
- H. Zhu, B. Chen, and C. Yang, Understanding why ViT trains badly on small datasets: An intuitive perspective, arXiv:2302.03751.
- Y. Zhang, A. Warstadt, H.-S. Li, and S. R. Bowman, When do you need billions of words of pretraining data? in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), edited by C. Zong, F. Xia, W. Li, and R. Navigli (Association for Computational Linguistics, 2021), pp. 1112–1125.
- T. W. Ko and S. P. Ong, Data-efficient construction of high-fidelity graph deep learning interatomic potentials, npj Comput. Mater. 11, 65 (2025).
- J. Kiechle, D. M. Lang, S. M. Fischer, L. Felsner, J. C. Peeken, and J. A. Schnabel, Graph neural networks: A suitable alternative to MLPs in latent 3D medical image classification? in Graphs in Biomedical Image Analysis (Springer, Cham, Switzerland, 2025), pp. 12–22.
- M. Oliva, S. Banik, J. Josifovski, and A. Knoll, Graph neural networks for relational inductive bias in vision-based deep reinforcement learning of robot control, in 2022 International Joint Conference on Neural Networks (IJCNN) (IEEE, 2022), pp. 1–9.
- J. H. Giraldo, K. Skianis, T. Bouwmans, and F. D. Malliaros, On the trade-off between over-smoothing and over-squashing in deep graph neural networks, in Proceedings of the 32nd ACM international Conference on Information and Knowledge Management (CIKM '23), Birmingham, United Kingdom (Association for Computing Machinery, New York, NY, 2023), pp. 566–576.
- T. K. Rusch, M. M. Bronstein, and S. Mishra, A survey on oversmoothing in graph neural networks, arXiv:2303.10993.
- Q. Yan, S. Kar, S. Chowdhury, and A. Bansil, The case for a defect genome initiative, Adv. Mater. 36, 2303098 (2024).
- E. Dramko, Y. Xiong, Y. Zhu, G. Hautier, T. Reps, C. Jermaine, and A. Kyrillidis, Dataset for ADAPT: Lightweight, long-range machine learning force fields without graphs [Dataset], Zenodo, 2025, https://doi.org/10.5281/zenodo.18962774.
- https://github.com/EvanDramko/ADAPT_Released.
- A. Jain, S. P. Ong, G. Hautier, W. Chen, W. D. Richards, S. Dacek, S. Cholia, D. Gunter, D. Skinner, and G. Ceder, Commentary: The Materials Project: A materials genome approach to accelerating materials innovation, APL Mater. 1, 011002 (2013).
- K. Mathew, J. H. Montoya, A. Faghaninia, S. Dwarakanath, M. Aykol, H. Tang, I. heng Chu, T. Smidt, B. Bocklund, M. Horton, J. Dagdelen, B. Wood, Z.-K. Liu, J. Neaton, S. P. Ong, K. Persson, and A. Jain, Atomate: A high-level interface to generate, execute, and analyze computational materials science workflows, Comput. Mater. Sci. 139, 140 (2017).
- S. P. Ong, W. D. Richards, A. Jain, G. Hautier, M. Kocher, S. Cholia, D. Gunter, V. L. Chevrier, K. A. Persson, and G. Ceder, Python materials genomics (pymatgen): A robust, open-source Python library for materials analysis, Comput. Mater. Sci. 68, 314 (2013).
- G. Kresse and J. Furthmüller, Efficient iterative schemes for ab initio total-energy calculations using a plane-wave basis set, Phys. Rev. B 54, 11169 (1996).
- G. Kresse and J. Furthmüller, Efficiency of ab-initio total energy calculations for metals and semiconductors using a plane-wave basis set, Comput. Mater. Sci. 6, 15 (1996).
- P. E. Blöchl, Projector augmented-wave method, Phys. Rev. B 50, 17953 (1994).
- C. Bergmeir, R. J. Hyndman, and B. Koo, A note on the validity of cross-validation for evaluating autoregressive time series prediction, Comput. Stat. Data Anal. 120, 70 (2018).
- L. Prechelt, Early stopping—But when? Neural Networks: Tricks of the Trade (Springer, 2002), pp. 55–69.
- M. Li, M. Soltanolkotabi, and S. Oymak, Gradient descent with early stopping is provably robust to label noise for overparameterized neural networks, arXiv:1903.11680.
- J. Deng, J. Guo, N. Xue, and S. Zafeiriou, ArcFace: Additive angular margin loss for deep face recognition, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2019), pp. 4690–4699.
- G. Ma, Y. Wang, D. Lim, S. Jegelka, and Y. Wang, A canonicalization perspective on invariant and equivariant learning, Adv. Neural Inf. Process. Syst. 37, 60936 (2024).
- J. Burgess, J. J. Nirschl, M.-C. Zanellati, A. Lozano, S. Cohen, and S. Yeung-Levy, Orientation-invariant autoencoders learn robust representations for shape profiling of cells and organelles, Nat. Commun. 15, 1022 (2024).
- C.-H. Chang, C.-N. Chou, and E. Y. Chang, Clkn: Cascaded Lucas-Kanade networks for image alignment, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2017), pp. 2213–2221.
- V. Boussot, C. Hémon, J.-C. Nunes, J. Dowling, S. Rouzé, C. Lafond, A. Barateau, and J.-L. Dillenseger, IMPACT: A generic semantic loss for multimodal medical image registration, arXiv:2503.24121.
- L. Pu, R. G. Govindaraj, J. M. Lemoine, H.-C. Wu, and M. Brylinski, DeepDrug3D: Classification of ligand-binding pockets in proteins with a convolutional neural network, PLoS Comput. Biol. 15, e1006718 (2019).
- J. Martinka, M. Pederzoli, M. Barbatti, P. O. Dral, and J. Pittner, A simple approach to rotationally invariant machine learning of a vector quantity, J. Chem. Phys. 161, 174104 (2024).
- C. R. Qi, H. Su, K. Mo, and L. J. Guibas, PointNet: Deep learning on point sets for 3D classification and segmentation, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2017), pp. 652–660.
- G. Chevrot, P. Calligari, K. Hinsen, and G. R. Kneller, Least constraint approach to the extraction of internal motions from molecular dynamics trajectories of flexible macromolecules, J. Chem. Phys. 135, 084110 (2011).