Development and Preliminary Validation of the ARROW-AI Framework for Responsible AI-Assisted Research Writing in Higher Education

Authors

  • Luijim Santos Jose
  • Reynaldo Cabual

Keywords:

academic integrity; generative artificial intelligence; higher education; research writing; responsible AI use

Abstract

Generative artificial intelligence is increasingly used in higher-education research writing, yet institutions often lack operational procedures to protect authorship, ensure research alignment, verify, and disclose. This study developed and preliminarily evaluated the ARROW-AI Framework and Toolkit through a sequential multiphase developmental-validation design with embedded mixed-methods formative evaluation and a bounded pilot. Participants comprised 12 experts, 84 student-researchers, 18 advisers/instructors, and 32 pilot cases. Descriptive expert-domain ratings ranged from 3.38 to 3.58 for the framework and from 3.20 to 3.53 for the toolkit, while all 15 components met the prespecified I-CVI and CVR criteria. Student-researcher domain ratings ranged from 4.05 to 4.18, and adviser/instructor ratings ranged from 3.82 to 4.24; documentation and workload burden were moderate. Pilot completion averaged 88.91%; six initially high-risk requests were down-scoped, and four cases required substantive revision. Because the instruments covered conceptually distinct formative domains, ratings were interpreted descriptively by domain rather than as scores from unidimensional psychometric scales. Integrated evidence supported six revisions, including shorter records, completed examples, clearer risk decisions, and stronger adviser guidance. ARROW-AI should therefore be treated as a developmental process guide rather than evidence of effectiveness or misconduct prevention.

https://doi.org/10.26803/ijlter.25.8.37

References

American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. https://www.testingstandards.net/open-access-files.html

Ayre, C., & Scally, A. J. (2014). Critical values for Lawshe’s content validity ratio: Revisiting the original methods of calculation. Measurement and Evaluation in Counseling and Development, 47(1), 79–86. https://doi.org/10.1177/0748175613513808

Bittle, K., & El-Gayar, O. (2025). Generative AI and academic integrity in higher education: A systematic review and research agenda. Information, 16(4), Article 296. https://doi.org/10.3390/info16040296

Chan, C. K. Y., & Hu, W. (2023). Students’ voices on generative AI: Perceptions, benefits, and challenges in higher education. International Journal of Educational Technology in Higher Education, 20(1), Article 43. https://doi.org/10.1186/s41239-023-00411-8

Chiu, T. K. F. (2024). Future research recommendations for transforming higher education with generative AI. Computers and Education: Artificial Intelligence, 6, Article 100197. https://doi.org/10.1016/j.caeai.2023.100197

Committee on Publication Ethics. (2023, February 13). Authorship and AI tools. https://publicationethics.org/cope-position-statements/ai-author

Cotton, D. R. E., Cotton, P. A., & Shipway, J. R. (2024). Chatting and cheating: Ensuring academic integrity in the era of ChatGPT. Innovations in Education and Teaching International, 61(2), 228–239. https://doi.org/10.1080/14703297.2023.2190148

Crompton, H., & Burke, D. (2023). Artificial intelligence in higher education: The state of the field. International Journal of Educational Technology in Higher Education, 20(1), Article 22. https://doi.org/10.1186/s41239-023-00392-8

DeVellis, R. F., & Thorpe, C. T. (2021). Scale development: Theory and applications (5th ed.). SAGE Publications.

Fetters, M. D., & Tajima, C. (2022). Joint displays of integrated data collection in mixed methods research. International Journal of Qualitative Methods, 21, 16094069221104564. https://doi.org/10.1177/16094069221104564

Gruenhagen, J. H., Sinclair, P. M., Carroll, J.-A., Baker, P. R. A., Wilson, A., & Demant, D. (2024). The rapid rise of generative AI and its implications for academic integrity: Students’ perceptions and use of chatbots for assistance with assessments. Computers and Education: Artificial Intelligence, 7, Article 100273. https://doi.org/10.1016/j.caeai.2024.100273

International Center for Academic Integrity. (2021). The fundamental values of academic integrity (3rd ed.). https://academicintegrity.org/images/pdfs/20019_ICAI-Fundamental-Values_R12.pdf

International Committee of Medical Journal Editors. (2026, January). Recommendations for the conduct, reporting, editing, and publication of scholarly work in medical journals. https://www.icmje.org/recommendations/

Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73. https://doi.org/10.1111/jedm.12000

Lawshe, C. H. (1975). A quantitative approach to content validity. Personnel Psychology, 28(4), 563–575. https://doi.org/10.1111/j.1744-6570.1975.tb01393.x

Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), Article 100779. https://doi.org/10.1016/j.patter.2023.100779

Mayring, P. (2022). Qualitative content analysis: A step-by-step guide. SAGE Publications. https://doi.org/10.4135/9781036231798

McKenney, S., & Reeves, T. C. (2018). Conducting educational design research (2nd ed.). Routledge. https://doi.org/10.4324/9781315105642

Miao, F., & Cukurova, M. (2024). AI competency framework for teachers. UNESCO. https://doi.org/10.54675/ZJTE2084

Miao, F., & Holmes, W. (2023). Guidance for generative AI in education and research. UNESCO. https://doi.org/10.54675/EWZM9535

Miao, F., Shiohira, K., & Lao, N. (2024). AI competency framework for students. UNESCO. https://doi.org/10.54675/JKJB9835

O’Sullivan, J., Lowry, C., Woods, R., & Conlon, T. (2025). Generative AI in higher education teaching and learning: National policy framework. Higher Education Authority. https://doi.org/10.82110/PX37-MP48

Perkins, M., Roe, J., & Furze, L. (2025). Reimagining the artificial intelligence assessment scale: A refined framework for educational assessment. Journal of University Teaching and Learning Practice, 22(7). https://doi.org/10.53761/rrm4y757

Perkins, M., Roe, J., Postma, D., McGaughran, J., & Hickerson, D. (2024). Detection of GPT-4 generated text in higher education: Combining academic judgement and software to identify generative AI tool misuse. Journal of Academic Ethics, 22(1), 89–113. https://doi.org/10.1007/s10805-023-09492-6

Pinho, I., Costa, A. P., & Pinho, C. (2025). Generative AI governance model in educational research. Frontiers in Education, 10, Article 1594343. https://doi.org/10.3389/feduc.2025.1594343

Proctor, E., Silmere, H., Raghavan, R., Hovmand, P., Aarons, G., Bunger, A., Griffey, R., & Hensley, M. (2011). Outcomes for implementation research: Conceptual distinctions, measurement challenges, and research agenda. Administration and Policy in Mental Health and Mental Health Services Research, 38(2), 65–76. https://doi.org/10.1007/s10488-010-0319-7

Rajabi, P., Taghipour, P., Cukierman, D., & Doleck, T. (2024). Unleashing ChatGPT’s impact in higher education: Student and faculty perspectives. Computers in Human Behavior: Artificial Humans, 2(2), Article 100090. https://doi.org/10.1016/j.chbah.2024.100090

Richey, R. C., & Klein, J. D. (2007). Design and development research: Methods, strategies, and issues. Routledge. https://doi.org/10.4324/9780203826034

Smith, S. M., Tate, M., Freeman, K., Walsh, A., Ballsun-Stanton, B., & Lane, M. (2026). A university framework for the responsible use of generative AI in research. Journal of Higher Education Policy and Management, 48(1), 17–36. https://doi.org/10.1080/1360080X.2025.2509187

Tertiary Education Quality and Standards Agency. (2025, September 24). Enacting assessment reform in a time of artificial intelligence. https://www.teqsa.gov.au/guides-resources/resources/corporate-publications/enacting-assessment-reform-time-artificial-intelligence

Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13(1), Article 14045. https://doi.org/10.1038/s41598-023-41032-5

Yusoff, M. S. B. (2019). ABC of content validation and content validity index calculation. Education in Medicine Journal, 11(2), 49–54. https://doi.org/10.21315/eimj2019.11.2.6

Downloads

Published

2026-08-30

How to Cite

Jose, L. S. ., & Cabual, R. . (2026). Development and Preliminary Validation of the ARROW-AI Framework for Responsible AI-Assisted Research Writing in Higher Education. International Journal of Learning, Teaching and Educational Research, 25(8), 946–967. Retrieved from https://ijlter.net/index.php/ijlter/article/view/3026