- Open Access
Introductory physics students’ valued features in AI feedback generated from self-crafted and engineered prompts
Phys. Rev. Phys. Educ. Res. 22, 010146 – Published 16 June, 2026
DOI: https://doi.org/10.1103/v3pj-8491
Abstract
Generative artificial intelligence (AI) has shown promising potential in education, particularly in providing real-time, customized, student-facing feedback in large-enrollment courses. A crucial skill that mediates effective use of AI is the systematic structuring of natural language instructions to AI models, commonly referred to as prompt engineering. This study has two objectives: (i) to investigate the sophistication of student-generated prompts when seeking feedback from AI and (ii) to examine trends associated with students’ preferences and perceptions of features in AI feedback generated from self-crafted prompts vs prompts incorporating prompt-engineering techniques and principles of effective feedback. Results indicate that student-generated prompts typically reflect only a subset of foundational prompt-engineering techniques. We also identified 12 distinct features that students attribute the usefulness of AI-generated feedback, spanning four broader themes: (i) evaluation, (ii) content, (iii) presentation, and (iv) depth. Finally, students overwhelmingly prefer feedback generated from structured prompts, particularly those combining prompt-engineering techniques with principles of effective feedback. Implications of these results—such as integrating the principles of effective feedback in design and delivery of feedback through AI systems, and incorporating prompt engineering in introductory physics courses—are discussed.
Physics Subject Headings (PhySH)
Article Text
References (75)
- J. Hattie and H. Timperley, The power of feedback, Rev. Educ. Res. 77, 81 (2007).
- D. Burke and J. Pieterick, Giving Students Effective Written Feedback (McGraw-Hill Education, London, 2010).
- R. Azevedo and R. M. Bernard, A meta-analysis of the effects of feedback in computer-based instruction, J. Educ. Comput. Res. 13, 111 (1995).
- D. Giamos, O. Doucet, and P.-M. Léger, Continuous performance feedback: Investigating the effects of feedback content and feedback sources on performance, motivation to improve performance and task engagement, J. Organ. Behav. Manage. 44, 194 (2024).
- D. J. Nicol and D. Macfarlane-Dick, Formative assessment and self-regulated learning: A model and seven principles of good feedback practice, Stud. Higher Educ. 31, 199 (2006).
- J. M. Reddig, A. Arora, and C. J. MacLellan, Generating in-context, personalized feedback for intelligent tutors with large language models, Int. J. Artif. Intell. Educ. 35, 1 (2025).
- T. Fan, J. Song, and Z. Guan, Integrating diagnostic assessment into curriculum: A theoretical framework and teaching practices, Lang. Test. Asia 11, 2 (2021).
- A. Sirnoorkar, A. P. Jambuge, K. D. Rainey, A. Adamson, B. R. Wilcox, and J. T. Laverty, Theoretical approach for providing feedback for instructors through a standardized assessment for undergraduate physics, in Proceedings of the 17th International Conference of the Learning Sciences-ICLS 2023 (International Society of the Learning Sciences, Montreal, 2023), pp. 1326–1329.
- J. T. Laverty, A. Sirnoorkar, A. P. Jambuge, K. D. Rainey, J. Weaver, A. Adamson, and B. R. Wilcox, A new paradigm for research-based assessment development, presented at PER Conf. 2022, Grand Rapids, MI, 10.1119/perc.2022.pr.Laverty.
- G. Polverini and B. Gregorcic, How understanding large language models can inform the use of ChatGPT in physics education, Eur. J. Phys. 45, 025701 (2024).
- G. Polverini and B. Gregorcic, Evaluating vision-capable chatbots in interpreting kinematics graphs: A comparative study of free and subscription-based models, Front. Educ. 9, 1452414 (2024).
- A. Sirnoorkar, D. Zollman, J. T. Laverty, A. J. Magana, N. S. Rebello, and L. A. Bryan, Student and AI responses to physics problems examined through the lenses of sensemaking and mechanistic reasoning, Comput. Educ. 7, 100318 (2024).
- A. Bralin, A. Sirnoorkar, Y. Zhang, and N. S. Rebello, Mapping the literature landscape of artificial intelligence and machine learning in physics education research, presented at PER Conf. 2024, Boston, MA, 10.1119/perc.2024.pr.Bralin.
- G. Kortemeyer, J. Nöhl, and D. Onishchuk, Grading assistance for a handwritten thermodynamics exam using artificial intelligence: An exploratory study, Phys. Rev. Phys. Educ. Res. 20, 020144 (2024).
- G. Kortemeyer, Performance of the pre-trained large language model GPT-4 on automated short answer grading, Discover Artif. Intell. 4, 47 (2024).
- G. Kortemeyer and J. Nöhl, Assessing confidence in AI-assisted grading of physics exams through psychometrics: An exploratory study, Phys. Rev. Phys. Educ. Res. 21, 010136 (2025).
- G.-G. Lee, E. Latif, X. Wu, N. Liu, and X. Zhai, Applying large language models and chain-of-thought for automatic scoring, Comput. Educ. 6, 100213 (2024).
- Y. Wei, R. Zhang, J. Zhang, D. Qi, and W. Cui, Research on intelligent grading of physics problems based on large language models, Educ. Sci. 15, 116 (2025).
- S. El-Adawy, A. MacDonagh, and M. Abdelhafez, Exploring large language models as formative feedback tools in physics, presented in PER Conf. 2024, Boston, MA, 10.1119/perc.2024.pr.El-Adawy.
- A. Meyer, T. Bleckmann, and G. Friege, Automatic feedback on physics tasks using open-source generative artificial intelligence, Int. J. Sci. Educ. 1 (2025).
- T. Wan and Z. Chen, Exploring generative AI assisted feedback writing for students’ written responses to a physics conceptual question with prompt engineering and few-shot learning, Phys. Rev. Phys. Educ. Res. 20, 010152 (2024).
- Z. Chen and T. Wan, Grading explanations of problem-solving process and generating feedback using large language models at human-level accuracy, Phys. Rev. Phys. Educ. Res. 21, 010126 (2025).
- E. Mills, A. Mizouri, and A. Peach, Prompting better feedback: A study of custom GPT for formative assessment in undergraduate physics, Educ. Sci. 15, 1058 (2025).
- W. Allen, A. Shanker, and N. S. Rebello, Students’ perceptions to a large language model’s generated feedback and scores of argumentation essays, arXiv:2508.14759.
- A. Almegren, H. S. Mahdi, A. N. Hazaea, J. K. Ali, and R. M. Almegren, Evaluating the quality of AI feedback: A comparative study of AI and human essay grading, Innov. Educ. Teach. Int. 62, 1858 (2025).
- A. Kinder, F. J. Briese, M. Jacobs, N. Dern, N. Glodny, S. Jacobs, and S. Leßmann, Effects of adaptive feedback generated by a large language model: A case study in teacher education, Comput. Educ. 8, 100349 (2025).
- C. Song and Y. Song, Enhancing academic writing skills and motivation: Assessing the efficacy of ChatGPT in AI-assisted language learning for EFL students, Front. Psychol. 14, 1260843 (2023).
- K. Guo, M. Pan, Y. Li, and C. Lai, Effects of an AI-supported approach to peer feedback on university EFL students’ feedback quality and writing ability, Internet Higher Educ. 63, 100962 (2024).
- E. Fokides and E. Peristeraki, Comparing ChatGPT’s correction and feedback comments with that of educators in the context of primary students’ short essays written in English and Greek, Educ. Inf. Technol. 30, 2577 (2025).
- C.-W. Kao and B. L. Reynolds, Timed second language writing performance: Effects of perceived teacher vs perceived automated feedback, Hum. Soc. Sci. Commun. 11, 1 (2024).
- T. Jansen, L. Höft, L. Bahr, J. Fleckenstein, J. Möller, O. Köller, and J. Meyer, Empirische Arbeit: Comparing generative AI and expert feedback to students’ writing: Insights from student teachers, Psychol. Erzieh. Unterr. 71, 80 (2024).
- J. Zare, F. Ranjbaran Madiseh, and A. Derakhshan, Generative AI and English essay writing: Exploring the role of ChatGPT in enhancing learners’ task engagement, Appl. Linguist. amaf045 (2025).
- S. T. S. Chan, N. P. K. Lo, and A. M. H. Wong, Enhancing university level English proficiency with generative AI: Empirical insights into automated feedback and learning outcomes, Contemp. Educ. Technol. 16, ep541 (2024).
- A. H. Alsofyani and A. M. Barzanji, The effects of ChatGPT-generated feedback on Saudi EFL learners’ writing skills and perception at the tertiary level: A mixed-methods study, J. Educ. Comput. Res. 63, 431 (2025).
- J. Escalante, A. Pack, and A. Barrett, AI-generated feedback on writing: Insights into efficacy and ENL student preference, Int. J. Educ. Technol. Higher Educ. 20, 57 (2023).
- S. Rüdian, J. Podelo, J. Kužílek, and N. Pinkwart, Feedback on feedback: Student’s perceptions for feedback from teachers and few-shot LLMs, in Proceedings of the 15th International Learning Analytics and Knowledge Conference, Dublin, Ireland (Association for Computing Machinery, New York, NY, 2025), pp. 82–92.
- A. Aljasser, Investigating EFL students’ perceptions of feedback: A comparative study of instructor and ChatGPT-generated responses in academic writing, Educ. Inf. Technol. 30, 22967 (2025).
- J. Steiss, T. Tate, S. Graham, J. Cruz, M. Hebert, J. Wang, Y. Moon, W. Tseng, M. Warschauer, and C. B. Olson, Comparing the quality of human and ChatGPT feedback of students’ writing, Learn. Instr. 91, 101894 (2024).
- K. ElEbyary and R. Shabara, ChatGPT-generated corrective feedback: Does it do what it says on the tin?, Teach. Engl. Technol. 24, 68 (2024).
- M. Abdi Tabari, A. Kushki, and Y. Wang, Comparing the effects of teacher-and AI-mediated corrective feedback on accuracy, complexity, and quality in L2 written narratives, Comput. Assisted Lang. Learn. 1 (2025).
- S. K. Banihashem, N. T. Kerman, O. Noroozi, J. Moon, and H. Drachsler, Feedback sources in essay writing: Peer-generated or AI-generated feedback?, Int. J. Educ. Technol. Higher Educ. 21, 23 (2024).
- S. K. Banihashem, O. Noroozi, H. Khosravi, C. D. Schunn, and H. Drachsler, Pedagogical framework for hybrid intelligent feedback, Innov. Educ. Teach. Int. 63, 1 (2025).
- Z. Zhang, S. Aubrey, X. Huang, and T. K. Chiu, The role of generative AI and hybrid feedback in improving L2 writing skills: A comparative study, Innov. Lang. Learn. Teach. 1 (2025).
- J. Han and M. Li, Exploring ChatGPT-supported teacher feedback in the EFL context, System 126, 103502 (2024).
- M. Navío-Inglés, J. Guzmán Mora, P. O’Connor-Jiménez, and A. García González, What’s next for feedback in writing instruction? Pre-service teachers’ perceptions of assessment practices and the role of generative AI, Educ. Sci. 15, 1534 (2025).
- F. Ghoochani, J. Scharfenberger, B. Funk, R. Doublan, M. Jakharabhai Odedra, and B. Etsiwah, From feedback to formative guidance: Leveraging LLMs for personalized support in programming projects, in Adjunct Proceedings of the 33rd ACM Conference on User Modeling, Adaptation and Personalization, New York City (Association for Computing Machinery, New York, NY, 2025), pp. 398–403.
- Ş. Çağlar-Özhan, P. Tekeli, and S. Arkün-Kocadere, Comparison of AI-generated and instructor feedback: No significant difference in perceived feedback quality and neither on performance, J. Comput. Assist. Learn. 41, e70134 (2025).
- E. Bauer, C. Richters, A. J. Pickal, M. Klippert, M. Sailer, and M. Stadler, Effects of AI-generated adaptive feedback on statistical skills and interest in statistics: A field experiment in higher education, Br. J. Educ. Technol. 56, 1735 (2025).
- Z. Zhang, Z. Dong, Y. Shi, T. Price, N. Matsuda, and D. Xu, Students’ perceptions and preferences of generative artificial intelligence feedback for programming, in Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence and Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence and Fourteenth Symposium on Educational Advances in Artificial Intelligence (AAAI Press, Vancouver, 2024), pp. 23250–23258.
- N. Ippisch, M. Herklotz, A.-C. Haensch, and C. Schwemmer, Beyond correctness: Evaluating and improving LLM feedback in statistical education, arXiv:2511.07628.
- S. Yeung, University students’ engagement with generative AI-supported automated writing evaluation (AWE) feedback, J. Second Lang. Writing 68, 101203 (2025).
- R. AlGhamdi, Exploring the impact of ChatGPT-generated feedback on technical writing skills of computing students: A blinded study, Educ. Inf. Technol. 29, 18901 (2024).
- Y. Zhou, M. Pankiewicz, L. Paquette, and R. S. Baker, Impact of LLM feedback on learner persistence in programming, in Proceedings of the 33rd International Conference on Computers in Education, ICCE 2025, Chennai, India (Asia-Pacific Society for Computers in Education, Chennai, 2025).
- L. V. Navarro, M. M. G. Pérez, P. Barajas-Á lvarez, E. Ortiz-Zavala, E. R. Villegas, and J. M. Sánchez-Lizárraga, Qualitative feedback comparison between professors and AI in STEM education, in Proceedings of the 2025 IEEE Global Engineering Education Conference (EDUCON), London, United Kingdom (IEEE, New York, 2025), pp. 1–6.
- T. Maillo, C. A. Alexandru, and A. Constantin, Enhancing marked with peer feedback functionality with LLM support, in Proceedings of the 2025 Conference on UK and Ireland Computing Education Research (2025), pp. 1–7.
- M. Steinbach, S. Bhandari, J. Meyer, and Z. A. Pardos, When LLMs hallucinate: Examining the effects of erroneous feedback in math tutoring systems, in Proceedings of the Twelfth ACM Conference on Learning Scale, Palermo Italy (Association for Computing Machinery, New York, NY, 2025), pp. 139–150.
- X. Dai, Z. Wen, J. Jiang, H. Liu, and Y. Zhang, How students use AI feedback matters: Experimental evidence on physics achievement and autonomy, arXiv:2505.08672.
- Q. Jia, J. Cui, R. Xi, C. Liu, P. Rashid, R. Li, and E. Gehringer, On assessing the faithfulness of LLM-generated feedback on student assignments, in Proceedings of the 17th International Conference on Educational Data Mining (International Educational Data Mining Society, Atlanta, 2024), pp. 491–499.
- A. Becerra and R. Cobos, Enhancing the professional development of engineering students through an AI-based collaborative feedback system, in Proceedings of the 2025 IEEE Global Engineering Education Conference (EDUCON), Global Engineering Education Conference (EDUCON) (IEEE, New York, 2025), pp. 1–9.
- L. W. Liebenow, F. T. Schmidt, J. Meyer, and J. Fleckenstein, Self-assessment accuracy in the age of artificial intelligence: Differential effects of LLM-generated feedback, Comput. Educ. 237, 105385 (2025).
- J. Weidlich, F. Gotsch, K. Schudel, C. Marusic-Würscher, J. Mazzarella, H. Bolten, D. Bütler, S. Luger, B. Wohlfehnder, and K. M. Merki, Teacher, peer, or AI? Comparing effects of feedback sources in higher education, Comput. Educ. Open 9, 100300 (2025).
- K. Seßler, A. Bewersdorff, C. Nerdel, and E. Kasneci, Towards adaptive feedback with AI: Comparing the feedback quality of LLMs and teachers on experimentation protocols, arXiv:2502.12842.
- L. J. Jacobsen and K. E. Weber, The promises and pitfalls of LLMs as feedback providers: A study of prompt engineering and the quality of AI-driven feedback, Cent. Open Sci. 10, 433 (2023).
- B. Chen, Z. Zhang, N. Langrené, and S. Zhu, Unleashing the potential of prompt engineering for large language models, Patterns 6, 101260 (2025).
- https://www.qualtrics.com/.
- R. W. Chabay and B. A. Sherwood, Matter and Interactions (John Wiley & Sons, New York, 2015).
- S. E. Toulmin, The Uses of Argument (Cambridge University Press, Cambridge, England, 2003).
- K. L. McNeill and J. Krajcik, Inquiry and scientific explanations: Helping students use evidence and reasoning, Science as Inquiry in the Secondary Setting, edited by J. Luft, R. L. Bell, and J. Gess-Newsome (2008), https://scholar.google.com/scholar?hl=en&as_sdt=0%2C5&q=McNeill+and+Krajick%2C+Inquiry+and+scientific+explanations%3A+Helping+students+use+evidence&btnG=.
- J. W. Creswell, W. E. Hanson, V. L. Clark Plano, and A. Morales, Qualitative research designs: Selection and implementation, Counselling psychology quarterly 35, 236 (2007).
- G. Sawalha, I. Taj, and A. Shoufan, Analyzing student prompts and their effect on ChatGPT’s performance, Cogent Educ. 11, 2397200 (2024).
- S. Tassoti, Assessment of students use of generative artificial intelligence: Prompting strategies and prompt engineering in chemistry education, J. Chem. Educ. 101, 2475 (2024).
- F. Valeri, P. Nilsson, and A.-M. Cederqvist, Exploring students’ experience of ChatGPT in STEM education, Comput. Educ. 8, 100360 (2025).
- M. Henderson, M. Bearman, J. Chung, T. Fawns, S. Buckingham Shum, K. E. Matthews, and J. de Mello Heredia, Comparing generative AI and teacher feedback: Student perceptions of usefulness and trustworthiness, Assess. Eval. Higher Educ. 1 (2025).
- S. Prompiengchai, C. Narreddy, and S. Joordens, A practical guide for supporting formative assessment and feedback using generative AI, arXiv:2505.23405.
- M. N. Dahlkemper, S. Z. Lahme, and P. Klein, How do physics students evaluate artificial intelligence responses on comprehension questions? A study on the perceived scientific accuracy and linguistic quality of ChatGPT, Phys. Rev. Phys. Educ. Res. 19, 010142 (2023).