Evaluating AI-Generated Physics Solutions: A Systematic Review of Solution Accuracy, Student Evaluation, and Implications for Critical Thinking in Higher Physics Education

Authors

  • Gladys Mahaut Institut Agama Islam Negeri Sorong

DOI:

https://doi.org/10.47945/create.v1i1.3355

Keywords:

Generative AI, Large Language Models, Physics Education, Critical Thinking, Learning From Errors, AI Literacy, Systematic Review

Abstract

Generative artificial intelligence (AI) is now capable of producing coherent, step-by-step solutions to physics problems. However, its implications for physics education are determined more by students’ ability to evaluate the validity of these solutions than by their availability. This systematic review, reported with reference to the PRISMA 2020 guidelines, synthesizes peer-reviewed physics education studies published between 2020 and 2026 to address three questions: (1) how accurately does generative AI solve or explain physics tasks, and what types of errors does it produce; (2) how do students evaluate AI-generated physics responses; and (3) what instructional opportunities and risks emerge for the development of critical thinking in higher physics education? Eleven studies were synthesized through deductive and inductive thematic analysis. AI accuracy was highly dependent on task type. Large language models performed at levels comparable to or above those of university students on several concept inventories and olympiad problems, but solved only 8.3% of underspecified problems (compared with 62.5% of well-specified problems), misinterpreted graphs and visual representations, and generated contradictory explanations. The dominant errors were related to physical modeling, assumptions, and representations rather than arithmetic. Students often failed to detect these errors: first- and second-year physics students (n = 102) rated incorrect AI responses to the most difficult problems as equivalent to correct worked solutions. Instructional studies reported positive attitudes and beneficial learning support, but none measured improvements in critical thinking using validated instruments. Thus, the claim that evaluating AI solutions develops critical thinking is theoretically plausible, grounded in research on learning from errors and critical-thinking skills, but has not yet been empirically demonstrated in physics education. This review proposes an Evaluate–Diagnose–Correct–Reflect framework and a research agenda that distinguishes procedural correctness from conceptual validity.

References

Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), Article e2422633122. https://doi.org/10.1073/pnas.2422633122

Bitzenbauer, P. (2023). ChatGPT in physics education: A pilot study on easy-to-implement activities. Contemporary Educational Technology, 15(3), Article ep430. https://doi.org/10.30935/cedtech/13176

Braun, V., & Clarke, V. (2006). Using thematic analysis in psychology. Qualitative Research in Psychology, 3(2), 77–101. https://doi.org/10.1191/1478088706qp063oa

Brown, D., & Cox, A. J. (2009). Innovative uses of video analysis. The Physics Teacher, 47(3), 145–150. https://pubs.aip.org/aapt/pte/article-abstract/47/3/145/275154/Innovative-Uses-of-Video-Analysis

Chi, M. T. H., Feltovich, P. J., & Glaser, R. (1981). Categorization and representation of physics problems by experts and novices. Cognitive Science, 5(2), 121–152. https://doi.org/10.1207/s15516709cog0502_2

Cotton, D. R. E., Cotton, P. A., & Shipway, J. R. (2024). Chatting and cheating: Ensuring academic integrity in the era of ChatGPT. Innovations in Education and Teaching International, 61(2), 228–239. https://doi.org/10.1080/14703297.2023.2190148

Crompton, H., & Burke, D. (2023). Artificial intelligence in higher education: The state of the field. International Journal of Educational Technology in Higher Education, 20, Article 22. https://doi.org/10.1186/s41239-023-00392-8

Dahlkemper, M. N., Lahme, S. Z., & Klein, P. (2023). How do physics students evaluate artificial intelligence responses on comprehension questions? A study on the perceived scientific accuracy and linguistic quality of ChatGPT. Physical Review Physics Education Research, 19(1), Article 010142. https://doi.org/10.1103/PhysRevPhysEducRes.19.010142

Docktor, J. L., & Mestre, J. P. (2014). Synthesis of discipline-based education research in physics. Physical Review Special Topics—Physics Education Research, 10(2), Article 020119. https://doi.org/10.1103/PhysRevSTPER.10.020119

Ennis, R. H. (1993). Critical thinking assessment. Theory Into Practice, 32(3), 179–186. https://doi.org/10.1080/00405849309543594

Facione, P. A. (1990). Critical thinking: A statement of expert consensus for purposes of educational assessment and instruction (ERIC Document No. ED315423). American Philosophical Association. https://eric.ed.gov/?id=ED315423

Gerlich, M. (2025). AI tools in society: Impacts on cognitive offloading and the future of critical thinking. Societies, 15(1), Article 6. https://doi.org/10.3390/soc15010006

Gregorcic, B., & Pendrill, A.-M. (2023). ChatGPT and the frustrated Socrates. Physics Education, 58(3), Article 035021. https://doi.org/10.1088/1361-6552/acc299

Große, C. S., & Renkl, A. (2007). Finding and fixing errors in worked examples: Can this foster learning outcomes? Learning and Instruction, 17(6), 612–634. https://doi.org/10.1016/j.learninstruc.2007.09.008

Halpern, D. F. (1998). Teaching critical thinking for transfer across domains: Dispositions, skills, structure training, and metacognitive monitoring. American Psychologist, 53(4), 449–455. https://doi.org/10.1037/0003-066X.53.4.449

Hong, Q. N., Fàbregues, S., Bartlett, G., Boardman, F., Cargo, M., Dagenais, P., Gagnon, M.-P., Griffiths, F., Nicolau, B., O’Cathain, A., Rousseau, M.-C., Vedel, I., & Pluye, P. (2018). The Mixed Methods Appraisal Tool (MMAT) version 2018 for information professionals and researchers. Education for Information, 34(4), 285–291. https://doi.org/10.3233/EFI-180221

Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), Article 248. https://doi.org/10.1145/3571730

Kasneci, E., Sessler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., Günnemann, S., Hüllermeier, E., Krusche, S., Kutyniok, G., Michaeli, T., Nerdel, C., Pfeffer, J., Poquet, O., Sailer, M., Schmidt, A., Seidel, T., … Kasneci, G. (2023). ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences, 103, Article 102274. https://doi.org/10.1016/j.lindif.2023.102274

Kohnke, L., Moorhouse, B. L., & Zou, D. (2023). ChatGPT for language teaching and learning. RELC Journal, 54(2), 537–550. https://doi.org/10.1177/00336882231162868

Kortemeyer, G. (2023). Could an artificial-intelligence agent pass an introductory physics course? Physical Review Physics Education Research, 19(1), Article 010132. https://doi.org/10.1103/PhysRevPhysEducRes.19.010132

Kregear, T., Babayeva, M., & Widenhorn, R. (2025). Analysis of student interactions with a large language model in an introductory physics lab setting. International Journal of Artificial Intelligence in Education, 35, 2993–3016. https://doi.org/10.1007/s40593-025-00489-3

Krupp, L., Steinert, S., Kiefer-Emmanouilidis, M., Avila, K. E., Lukowicz, P., Kuhn, J., Küchemann, S., & Karolus, J. (2024). Unreflected acceptance—Investigating the negative consequences of ChatGPT-assisted problem solving in physics education. In HHAI 2024: Hybrid human AI systems for the social good (Frontiers in Artificial Intelligence and Applications). IOS Press. https://doi.org/10.3233/FAIA240195

Küchemann, S., Steinert, S., Revenga, N., Schweinberger, M., Dinc, Y., Avila, K. E., & Kuhn, J. (2023). Can ChatGPT support prospective teachers in physics task development? Physical Review Physics Education Research, 19(2), Article 020128. https://doi.org/10.1103/PhysRevPhysEducRes.19.020128

Kuo, E., Hull, M. M., Gupta, A., & Elby, A. (2013). How students blend conceptual and formal mathematical reasoning in solving physics problems. Science Education, 97(1), 32–57. https://doi.org/10.1002/sce.21043

Lo, C. K. (2023). What is the impact of ChatGPT on education? A rapid review of the literature. Education Sciences, 13(4), Article 410. https://doi.org/10.3390/educsci13040410

Lubis, H., Van Harling, V. N., & Panunggul, V. B. (2025). Enhancing critical thinking in physics education through AI: A systematic literature review of trends and pedagogical implications. Jurnal Pendidikan dan Ilmu Fisika, 5(2), 343–354. https://doi.org/10.52434/jpif.v5i2.43353

Metcalfe, J. (2017). Learning from errors. Annual Review of Psychology, 68, 465–489. https://doi.org/10.1146/annurev-psych-010416-044022

Ng, D. T. K., Leung, J. K. L., Chu, S. K. W., & Qiao, M. S. (2021). Conceptualizing AI literacy: An exploratory review. Computers and Education: Artificial Intelligence, 2, Article 100041. https://doi.org/10.1016/j.caeai.2021.100041

OpenAI. (2023). GPT-4 technical report (arXiv:2303.08774). arXiv. https://doi.org/10.48550/arXiv.2303.08774

Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., Shamseer, L., Tetzlaff, J. M., Akl, E. A., Brennan, S. E., Chou, R., Glanville, J., Grimshaw, J. M., Hróbjartsson, A., Lalu, M. M., Li, T., Loder, E. W., Mayo-Wilson, E., McDonald, S., … Moher, D. (2021). The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ, 372, Article n71. https://doi.org/10.1136/bmj.n71

Parasuraman, R., & Manzey, D. H. (2010). Complacency and bias in human use of automation: An attentional integration. Human Factors, 52(3), 381–410. https://doi.org/10.1177/0018720810376055

Polverini, G., & Gregorcic, B. (2024a). How understanding large language models can inform the use of ChatGPT in physics education. European Journal of Physics, 45(2), Article 025701. https://doi.org/10.1088/1361-6404/ad1420

Polverini, G., & Gregorcic, B. (2024b). Performance of ChatGPT on the test of understanding graphs in kinematics. Physical Review Physics Education Research, 20(1), Article 010109. https://doi.org/10.1103/PhysRevPhysEducRes.20.010109

Polverini, G., Melin, J., Önerud, E., & Gregorcic, B. (2025). Performance of ChatGPT on tasks involving physics visual representations: The case of the brief electricity and magnetism assessment. Physical Review Physics Education Research, 21(1), Article 010154. https://doi.org/10.1103/PhysRevPhysEducRes.21.010154

Risko, E. F., & Gilbert, S. J. (2016). Cognitive offloading. Trends in Cognitive Sciences, 20(9), 676–688. https://doi.org/10.1016/j.tics.2016.07.002

Sallam, M. (2023). ChatGPT utility in healthcare education, research, and practice: Systematic review on the promising perspectives and valid concerns. Healthcare, 11(6), Article 887. https://doi.org/10.3390/healthcare11060887

Tlili, A., Shehata, B., Adarkwah, M. A., Bozkurt, A., Hickey, D. T., Huang, R., & Agyemang, B. (2023). What if the devil is my guardian angel: ChatGPT as a case study of using chatbots in education. Smart Learning Environments, 10, Article 15. https://doi.org/10.1186/s40561-023-00237-x

Tschisgale, P., Maus, H., Kieser, F., Kroehs, B., Petersen, S., & Wulff, P. (2025). Evaluating GPT- and reasoning-based large language models on Physics Olympiad problems: Surpassing human performance and implications for educational assessment. Physical Review Physics Education Research. https://doi.org/10.1103/6fmx-bsnl

UNESCO. (2024). AI competency framework for students. UNESCO. https://www.unesco.org/en/articles/ai-competency-framework-students

VanLehn, K. (2011). The relative effectiveness of human tutoring, intelligent tutoring systems, and other tutoring systems. Educational Psychologist, 46(4), 197–221. https://doi.org/10.1080/00461520.2011.611369

Wang, K. D., Burkholder, E., Wieman, C., Salehi, S., & Haber, N. (2024). Examining the potential and pitfalls of ChatGPT in science and engineering problem-solving. Frontiers in Education, 8, Article 1330486. https://doi.org/10.3389/feduc.2023.1330486

Zawacki-Richter, O., Marín, V. I., Bond, M., & Gouverneur, F. (2019). Systematic review of research on artificial intelligence applications in higher education – where are the educators? International Journal of Educational Technology in Higher Education, 16, Article 39. https://doi.org/10.1186/s41239-019-0171-0

Downloads

Published

2026-05-25

How to Cite

Mahaut, G. (2026). Evaluating AI-Generated Physics Solutions: A Systematic Review of Solution Accuracy, Student Evaluation, and Implications for Critical Thinking in Higher Physics Education. CREATE: Contemporary Research in Education and Teaching, 1(1), 32–42. https://doi.org/10.47945/create.v1i1.3355