AI-driven Assessment in Education: A Bibliometric Analysis and Framework for Assessment Design and Governance
Keywords:
artificial intelligence; generative AI; educational assessment; bibliometric analysis; human-AI collaboration; assessment governance; educationAbstract
The rapid advancement of artificial intelligence (AI), particularly generative AI and large language models, has transformed educational assessment practices, raising concerns regarding assessment validity, academic integrity, human judgement, and governance. This study aims to map the intellectual landscape of AI-driven assessment research and synthesise key developments that inform the design, implementation, and governance of AI-supported assessment. A bibliometric analysis was conducted on 820 Scopus-indexed publications published between 2015 and 2025. Using VOSviewer, publication trend analysis, keyword co-occurrence analysis, and bibliographic coupling analysis were employed to examine research productivity, thematic evolution, and intellectual structures within the field. The findings reveal a substantial increase in publications following the emergence of generative AI technologies, particularly between 2023 and 2025. Four major thematic areas were identified: generative AI and educational innovation, assessment and evaluation, AI technologies and learning systems, and disciplinary applications. Bibliographic coupling analyses further indicate growing intellectual convergence around assessment redesign, human-AI collaboration, academic integrity, and responsible AI use. These findings suggest that AI-driven assessment has evolved beyond technological implementation toward a broader socio-technical discourse encompassing pedagogical, ethical, and governance considerations. Based on the synthesis of the findings, this study proposes the Human-AI Governance and Design (HAGD) Framework, comprising three interdependent dimensions: Assessment Design, Human-AI Role Allocation, and Governance and Oversight. The framework provides a structured foundation for the responsible integration of AI into educational assessment while maintaining educational quality, fairness, transparency, and public trust. The study contributes to the growing discourse on AI-driven assessment and offers practical insights for educators, assessment designers, and policymakers.
https://doi.org/10.26803/ijlter.25.7.31
References
Adams, J. (2013). The fourth age of research. Nature, 497(7451), 557-560. https://doi.org/10.1038/497557a
Awidi, I. T. (2024). Comparing expert tutor evaluation of reflective essays with marking by generative artificial intelligence (AI) tool. Computers and Education: Artificial Intelligence, 6, 100226. https://doi.org/10.1016/j.caeai.2024.100226
Boyack, K. W., & Klavans, R. (2010). Co-citation analysis, bibliographic coupling, and direct citation: Which citation approach represents the research front most accurately? Journal of the American Society for Information Science and Technology, 61(12), 2389-2404. https://doi.org/10.1002/asi.21419
Chen, L., Chen, P., & Lin, Z. (2020). Artificial intelligence in education: A review. IEEE Access, 8, 75264-75278. https://doi.org/10.1109/ACCESS.2020.2988510
Cope, B., Kalantzis, M., & Searsmith, D. (2021). Artificial intelligence for education: Knowledge and its assessment in AI-enabled learning ecologies. Educational Philosophy and Theory, 53(12), 1229-1245. https://doi.org/10.1080/00131857.2020.1728732
Fagbohun, O., Iduwe, N. P., Abdullahi, M., Ifaturoti, A., & Nwanna, O. M. (2024). Beyond traditional assessment: Exploring the impact of large language models on grading practices. Journal of Artificial Intelligence, Machine Learning and Data Science, 2(1), 1-8. https://doi.org/10.51219/JAIMLD/oluwole-fagbohun/19
Flodén, J. (2025). Grading exams using large language models: A comparison between human and AI grading of exams in higher education using ChatGPT. British Educational Research Journal, 51, 201–224. https://doi.org/10.1002/berj.4069
Gazni, A., & Didegah, F. (2016). The relationship between authors’ bibliographic coupling and citation exchange: Analyzing disciplinary differences. Scientometrics, 107(2), 609-626. https://doi.org/10.1007/s11192-016-1856-y
Glänzel, W., & Czerwon, H. J. (1995). A new methodological approach to bibliographic coupling and its application to research-front and other core documents. In Proceedings of the Fifth International Conference of the International Society for Scientometrics and Informetrics (pp. 167-176).
Glänzel, W., & Schubert, A. (2004). Analysing scientific networks through co-authorship. In H. F. Moed, W. Glänzel, & U. Schmoch (Eds.), Handbook of quantitative science and technology research (pp. 257-276). Springer. https://doi.org/10.1007/1-4020-2755-9_12
Giannakos, M., Azevedo, R., Brusilovsky, P., Cukurova, M., Dimitriadis, Y., Hernandez-Leo, D., ... & Rienties, B. (2025). The promise and challenges of generative AI in education. Behaviour & Information Technology, 44(11), 2518-2544. https://doi.org/10.1080/0144929X.2024.2394886
Halkiopoulos, C., & Gkintoni, E. (2024). Leveraging AI in e-learning: Personalized learning and adaptive assessment through cognitive neuropsychology - A systematic analysis. Electronics, 13(18), 3762. https://doi.org/10.3390/electronics13183762
Hamid, H., Zulkifli, K., Naimat, F., Che Yaacob, N. L., & Ng, K. W. (2023). Exploratory study on student perception on the use of chat AI in process-driven problem-based learning. Currents in Pharmacy Teaching and Learning, 15(12), 1017-1025. https://doi.org/10.1016/j.cptl.2023.10.001
Harzing, A. W., & Alakangas, S. (2016). Google Scholar, Scopus and the Web of Science: A longitudinal and cross-disciplinary comparison. Scientometrics, 106(2), 787-804. https://doi.org/10.1007/s11192-015-1798-9
Hooda, M., Rana, C., Dahiya, O., Rizwan, A., & Hossain, M. S. (2022). Artificial intelligence for assessment and feedback to enhance student success in higher education. Mathematical Problems in Engineering, 2022, Article 5215722. https://doi.org/10.1155/2022/5215722
Holmes, W., Bialik, M., & Fadel, C. (2019). Artificial intelligence in education: Promises and implications for teaching and learning. Center for Curriculum Redesign. https://curriculumredesign.org/wp-content/uploads/AIED-Book-Excerpt-CCR.pdf
Imran, M., & Almusharraf, N. (2024). Google Gemini as a next generation AI educational tool: A review of emerging educational technology. Smart Learning Environments, 11(1), 22. https://doi.org/10.1186/s40561-024-00310-z
Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1-73. https://doi.org/10.1111/jedm.12000
Jarneving, B. (2007). Bibliographic coupling and its application to research-front and other core documents. Journal of Informetrics, 1(4), 287-307. https://doi.org/10.1016/j.joi.2007.07.004
Jukiewicz, M. (2024). The future of grading programming assignments in education: The role of ChatGPT in automating the assessment and feedback process. Thinking Skills and Creativity, 52, 101522. https://doi.org/10.1016/j.tsc.2024.101522
Khlaif, Z. N., Ayyoub, A., Hamamra, B., Bensalem, E., Mitwally, M. A. A., Hattab, M. K., & Shadid, F. (2024). University teachers’ views on the adoption and integration of generative AI tools for student assessment in higher education. Education Sciences, 14(10), 1090. https://doi.org/10.3390/educsci14101090
Klarin, A. (2024). How to conduct a bibliometric content analysis: Guidelines and contributions of content co-occurrence or co-word literature reviews. International Journal of Consumer Studies, 48(2), e13031. https://doi.org/10.1111/ijcs.13031
Lelescu, A., & Kabiraj, S. (2024). Digital assessment in higher education: Sustainable trends and emerging frontiers in the AI era. In G. Grosseck, S. Sava, G. Ion, & L. M?lita (Eds.), Digital assessment in higher education (Lecture Notes in Educational Technology). Springer. https://doi.org/10.1007/978-981-97-6136-4_2
Leydesdorff, L., & Wagner, C. (2008). International collaboration in science and the formation of a core group. Journal of Informetrics, 2(4), 317-325. https://doi.org/10.1016/j.joi.2008.07.003
Liu, R. L. (2017). A new bibliographic coupling measure with descriptive capability. Scientometrics, 110(2), 919-935. https://doi.org/10.1007/s11192-016-2200-2
Luckin, R., Holmes, W., Griffiths, M., & Forcier, L. B. (2016). Intelligence unleashed: An argument for AI in education. Pearson.
Lye, C. Y., & Lim, L. (2024). Generative artificial intelligence in tertiary education: Assessment redesign principles and considerations. Education Sciences, 14(6), 569. https://doi.org/10.3390/educsci14060569
Mao, J., Chen, B., & Liu, J. C. (2024). Generative artificial intelligence in education and its implications for assessment. TechTrends, 68, 58-66. https://doi.org/10.1007/s11528-023-00911-4
Messick, S. (1989). Validity. In R. L. Linn (Ed.), Educational measurement (3rd ed., pp. 13–103). American Council on Education and Macmillan.
Molenaar, I. (2022). The concept of hybrid human-AI regulation: Exemplifying how to support young learners’ self-regulated learning. Computers and Education: Artificial Intelligence, 3, 100070. https://doi.org/10.1016/j.caeai.2022.100070
Mohamad Hanefar, S.B., Benaouda, B., Faizuddin, A. et al. Mapping the Landscape of Spiritual Intelligence: A Bibliometric Analysis of Trends, Patterns and Future Directions. Journal of Religion and Health, 64, 3419-3447 (2025). https://doi.org/10.1007/s10943-025-02386-4
Naidu, K., & Sevnarayan, K. (2023). ChatGPT: An ever-increasing encroachment of artificial intelligence in online assessment in distance education. Online Journal of Communication and Media Technologies, 13(3), e202336. https://doi.org/10.30935/ojcmt/13291
Nikolic, S., Daniel, S., Haque, R., Neal, P., & Sandison, C. (2023). ChatGPT versus engineering education assessment: A multidisciplinary and multi-institutional benchmarking and analysis of this generative artificial intelligence tool to investigate assessment integrity. European Journal of Engineering Education, 48(4), 559-614. https://doi.org/10.1080/03043797.2023.2213169
Nikolic, S., Sandison, C., Haque, R., Hassan, G. M., & Neal, P. (2024). ChatGPT, Copilot, Gemini, SciSpace and Wolfram versus higher education assessments: An updated multi-institutional study of the academic integrity impacts of generative artificial intelligence (GenAI) on assessment, teaching and learning in engineering. Australasian Journal of Engineering Education, 29(2), 126-153. https://doi.org/10.1080/22054952.2024.2372154
Oc, Y., Gonsalves, C., & Quamina, L. T. (2025). Generative AI in higher education assessments: Examining risk and tech-Savviness on students’ adoption. Journal of Marketing Education, 47(2), 138-155. https://doi.org/10.1177/02734753241302459
Prasad, R. D., Lim, S. P., Che Yob, F. S., Wong, Y. V., Magulod, G. C., Jr., & Adom, D. (2025). Navigating the tech turn: A bibliometric analysis of decision-making trends in 21st-century education. International Journal of Learning, Teaching and Educational Research, 24(11), 297–313. https://doi.org/10.26803/ijlter.24.11.14
Ray, P. P. (2023). ChatGPT: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope. Internet of Things and Cyber-Physical Systems, 3, 121-154. https://doi.org/10.1016/j.iotcps.2023.04.003
Roca, M. D. L., Chan, M. M., Garcia-Cabot, A., Garcia-Lopez, E., & Amado-Salvatierra, H. (2024). The impact of a chatbot working as an assistant in a course for supporting student learning and engagement. Computer Applications in Engineering Education, 32(5), e22750. https://doi.org/10.1002/cae.22750
Rudolph, J., Tan, S., & Tan, S. (2023). ChatGPT: Bullshit spewer or the end of traditional assessments in higher education? Journal of Applied Learning and Teaching, 6(1), 342-363. https://doi.org/10.37074/jalt.2023.6.1.9
Smolansky, A., Cram, A., Raduescu, C., Zeivots, S., Huber, E., & Kizilcec, R. F. (2023). Educator and student perspectives on the impact of generative AI on assessments in higher education. In Proceedings of the Tenth ACM Conference on Learning @ Scale (L@S '23) (pp. 378-382). Association for Computing Machinery. https://doi.org/10.1145/3573051.3596191
Sweeney, S. (2023). Who wrote this? Essay mills and assessment – Considerations regarding contract cheating and AI in higher education. International Journal of Management Education, 21(2), 100818. https://doi.org/10.1016/j.ijme.2023.100818
Tossell, C. C., Tenhundfeld, N. L., Momen, A., Cooley, K., & De Visser, E. J. (2024). Student perceptions of ChatGPT use in a college essay assignment: Implications for learning, grading, and trust in artificial intelligence. IEEE Transactions on Learning Technologies, 17, 1069-1081.
Usher, M. (2025). Generative AI vs instructor vs peer assessments: A comparison of grading and feedback in higher education. Assessment & Evaluation in Higher Education, 50(6), 1-16. https://doi.org/10.1080/02602938.2025.2487495
van Eck, N. J., & Waltman, L. (2010). Software survey: VOSviewer, a computer program for bibliometric mapping. Scientometrics, 84(2), 523-538. https://doi.org/10.1007/s11192-009-0146-3
Vashishth, T. K., Sharma, V., Sharma, K. K., Kumar, B., Panwar, R., & Chaudhary, S. (2024). AI-driven learning analytics for personalized feedback and assessment in higher education. In T. V. T. Nguyen & N. T. M. Vo (Eds.), Using traditional design methods to enhance AI-driven decision making (pp. 206-230). IGI Global. https://doi.org/10.4018/979-8-3693-0639-0.ch009
Villegas-Ch, W. E., Govea, J., Gutierrez, R., & Mera-Navarrete, A. (2024). Improving interaction and assessment in hybrid educational environments: An integrated approach in Microsoft Teams with the use of AI techniques. IEEE Access, 12, 93723-93738.
Vittorini, P., Menini, S., & Tonelli, S. (2021). An AI-based system for formative and summative assessment in data science courses. International Journal of Artificial Intelligence in Education, 31, 159-185. https://doi.org/10.1007/s40593-020-00230-2
Wagner, C. S., Whetsell, T. A., & Leydesdorff, L. (2017). Growth of international collaboration in science: Revisiting six specialties. Scientometrics, 110(3), 1633-1652. https://doi.org/10.1007/s11192-016-2230-9
Xiao, Y., & Watson, M. (2019). Guidance on conducting a systematic literature review. Journal of Planning Education and Research, 39(1), 93-112. https://doi.org/10.1177/0739456X17723971
Xu, H., Gan, W., Qi, Z., Wu, J., & Yu, P. S. (2024). Large language models for education: A survey. arXiv. https://doi.org/10.48550/arXiv.2405.13001
Zawacki-Richter, O., Marín, V. I., Bond, M., & Gouverneur, F. (2019). Systematic review of research on artificial intelligence applications in higher education – Where are the educators? International Journal of Educational Technology in Higher Education, 16(39), 1-27. https://doi.org/10.1186/s41239-019-0171-0
Zhai, X., & Nehm, R. H. (2023). AI and formative assessment: The train has left the station. Journal of Research in Science Teaching, 60(6), 1390-1398. https://doi.org/10.1002/tea.21885
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Shamsiah Banu Mohamad Hanefar, Maryam Ikram, Mutia Sobihah Abd Halim, Farah Naaz Abd Yunos, Anika Rahman

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
All articles published by IJLTER are licensed under a Creative Commons Attribution Non-Commercial No-Derivatives 4.0 International License (CCBY-NC-ND4.0).