CAN MACHINES WRITE YOUR EXAMS? A SYSTEMATIC REVIEW OF GENERATIVE ARTIFICIAL INTELLIGENCE FOR MULTIPLE-CHOICE QUESTION DEVELOPMENT IN MEDICAL EDUCATION
Date
2026-05-12
Authors
Ahmed M.J. Rabee
Abdulrahman Kharsa
Ammar Y.M. Saadawy
Mohammad Wasim Alnajjar
Ahmad M.A. Adi
Majid Ali
Journal Title
Journal ISSN
Volume Title
Publisher
Deanship of Scientific Research - Al-Quds University
Abstract
Background: Multiple-choice questions (MCQs) are essential for assessing medical education, whereas conventional item development is labour-intensive, inconsistent in quality, and often mismatched with evolving curricula. Generative artificial intelligence (AI), particularly large language models such as ChatGPT, has emerged as a promising tool for streamlining MCQ creation. However, evidence on the quality, psychometric properties, and educational effectiveness of AI-generated MCQs remains fragmented. Study Purpose/
Objectives: This review aimed to systematically assess the quality and effectiveness of generative AI in producing MCQs for medical education. Examining psychometric properties, content quality, cognitive level, resource efficiency, and user perceptions compared to human-authored questions.
Methods: This systematic review followed PRISMA 2020 guidelines and was registered on PROSPERO (CRD420251153109). Six databases (PubMed, Scopus, Embase, Web of Science, ERIC, and Europe PMC) were searched from January 2018 to November 2025. Two reviewers independently screened studies, extracted data, and assessed methodological quality using modified MERSQI and the Newcastle-Ottawa Scale.