Generative AI and item writing

A collection of resources on generative AI and its applications to producing language test items


Resources by Olena Rossi

Rossi, O., & Montcada, J. M. (2026). Beyond perfecting the prompt: Can text-mapping enhance the quality of AI-generated B2 MC reading items? Presentation delivered at the EALTA annual conference. Siena, Italy. June 2026. Download slides

Rossi, O, & Montcada, J. M. (2025). Generating multiple-choice items for a B2 English reading test with GPT-4: Targeting higher-order cognitive processing. Preprint. https://doi.org/10.35542/osf.io/qjfpb_v1 [free download]

Rossi, O., & Montcada, J. M. (2025). Generating reading test items with AI: Targeting higher-order thinking skills. Presentation delivered at the annual EALTA conference. May 2025, Salzburg. Download slides

Rossi, O., & Montcada, J. M. (2025).  Using ChatGPT to generate True/False reading comprehension items: Recommendations for practice. Presentation delivered at the EALTA’s AI SIG online conference. March 2025. Download slides Watch presentation

Rossi, O. (2024). Automated item generation: An item writer perspective. Talk delivered at the BAAL TEASIG webinar. December 2024, online. Download slides

Rossi, O. (2024). Using ChatGPT to generate tasks for EAP reading and listening assessments. Workshop delivered as part of BALEAP Assessment Roadshow. July 2024, online. Download slides Watch workshop Part 1 Watch workshop Part 2

Rossi, O. (2024). Item writing with generative AI: Current issues and future directions. Presentation delivered at the inaugural meeting of the EALTA SIG Artificial Intelligence for Language Assessment. June 2024, Belfast. Download slides

Rossi, O. (2024). Assessment of language through AI: Opportunities, challenges, and future directions. Plenary talk delivered at 2024 International Conference Language Education 4.0: A Paradigm Shift towards Action-Oriented Approach, Artificial Intelligence Integration and Beyond. June 2024, Ankara Download slides

Rossi, O. (2023). Using technology to write language test items. Talk delivered at IATEFL TEASIG online conference Developing Assessment Tasks for the Classroom. September 2023. Download slides

Rossi, O. (2023). Using AI for test item generation: Opportunities and challenges. Webinar delivered as part of the EALTA Webinar series. May 2023. Download workshop slides Watch webinar 


Review studies (click the arrow to view the abstract)

Circi, R., Hicks, J., & Sikali, E. (2023). Automatic item generation: Foundations and machine learning-based approaches for assessments. Frontiers in Education, 8:858273.  https://doi.org/10.3389/feduc.2023.858273

This mini review summarizes the current state of knowledge about automatic item generation in the context of educational assessment and discusses key points in the item generation pipeline. Assessment is critical in all learning systems and digitalized assessments have shown significant growth over the last decade. This leads to an urgent need to generate more items in a fast and efficient manner. Continuous improvements in computational power and advancements in methodological approaches, specifically in the field of natural language processing, provide new opportunities as well as new challenges in automatic generation of items for educational assessment. This mini review asserts the need for more work across a wide variety of areas for the scaled implementation of AIG.

Song, Y., Du, J., & Zheng, Q. (2025). Automatic Item Generation for Educational Assessments: A Systematic Literature Review. Interactive Learning Environments. https://doi.org/10.1080/10494820.2025.2482588

This study reviewed automatic item generation (AIG) applications for educational assessments from 2010 to 2024. The analysis included 71 articles and focused on examining types of generated items and assessments, technical approaches, and evaluation models. The results showed that most generated items related to multiple choice questions, and the generated assessments were mainly about computer and medical sciences at college and vocational levels. The technical approaches were classified into four categories: feature engineering, architecture engineering, objective engineering, and prompt engineering. The models employed for evaluation were defined as manual annotation, man-machine collaborative evaluation, item analysis, Turing test, and value-added models. These findings provided knowledge and understanding to researchers and practitioners, showing the significance of expanding research focus, maintaining the theoretical foundation about educational assessments, and enhancing evaluation evidence for future AIG research.

Tan, B., Armoush, N., Mazzullo, E., Bulut, O., & Gierl, M. (2025). A Review of Automatic Item Generation Techniques Leveraging Large Language Models. International Journal of AIG & Testing Education, 12(2), 317–340.  https://doi.org/10.21449/ijate.1602294

This study reviews existing research on the use of large language models (LLMs) for automatic item generation (AIG). We performed a comprehensive literature search across seven research databases, selected studies based on predefined criteria, and summarized 60 relevant studies that employed LLMs in the AIG process. We identified the most commonly used LLMs in current AIG literature, their specific applications in the AIG process, and the characteristics of the generated items. We found that LLMs are flexible and effective in generating various types of items across different languages and subject domains. However, many studies have overlooked the quality of the generated items, indicating a lack of a solid educational foundation. Therefore, we share two suggestions to enhance the educational foundation for leveraging LLMs in AIG, advocating for interdisciplinary collaborations to exploit the utility and potential of LLMs.

Latest research (click the arrow to view the abstract)

Aryadoust, V., & Wong, J. (2026). How to train your dragon: Evaluating prompting and fine-tuning for GPT-based item generation in L2 listening assessment. Computers and Education: Artificial Intelligence. Advance online publication. https://doi.org/10.1016/j.caeai.2026.100623

Recent advances in large language models (LLMs), notably GPT models, have introduced new possibilities for automatic item generation (AIG) through prompt engineering. Although iterative prompt refinement can yield measurable improvements, this process eventually plateaus, as outputs remain inconsistent or misaligned with assessment constructs. This study reports an experimental comparison of prompting and fine-tuning to advance AIG for L2 listening assessment. First, we employed prompting and refined the instruction design over three successive iterations to determine an optimized prompt. We then fine-tuned GPT-4.1 using the same optimized prompt, holding prompt design constant to isolate the effect of model  adaptation. We generated a total of 40 tests and 240 multiple-choice items in the four model conditions: three prompt-only iterations and one fine-tuned iteration. To contextualize model performance against professional assessment standards, we also compared the GPT-generated items with 245 expert-generated listening items used as the fine-tuning dataset. The generated items were evaluated using a hybrid two-tiered framework that combined rule-based metrics with human review to assess content validity and scalability. Fine-tuning produced stronger outcomes overall, yielding items that were more contextually grounded, linguistically coherent, and balanced than those generated through prompting alone, with some requiring minimal or no revision; however, generating higher-order items involving discourse-level reasoning remained challenging. The expert comparison showed that fine-tuned items performed comparably to expert-authored items on passage dependence but remained relatively weaker in avoiding absolute language and in targeting localized spans of necessary information. Additionally, issues such as longest-correct-option bias and uneven key distribution persisted, indicating limitations inherent in LLM-generated items. These findings demonstrate the value of fine-tuning for improving item quality and stability, while underscoring the continued need for multidimensional evaluation frameworks and expert benchmarking to ensure construct-aligned, valid assessment design.

Flor, M., Wang, Z., Deane, P., & O’Reilly, T. (2026). Toward an automatic method for generating topical vocabulary test forms for specific reading passages (Research Report No. RR-26-02). ETS. https://doi.org/10.64634/kb8bf328

Background knowledge is typically needed for successful comprehension of topical and domain-specific reading passages, such as in the STEM domains. However, there are few automated measures of student knowledge that can be readily deployed and scored in time to make predictions on whether a given student will likely be able to understand a specific content-area text. In this research report, we present our effort in developing the K-tool, an automated system for generating topical vocabulary tests that measure students’ background knowledge related to a specific text. The system automatically detects the topic of a given text and produces topical vocabulary items based on their relationship with the topic. This information is used to automatically generate background knowledge forms that contain words that are highly related to the topic and share similar features but do not share high associations to the topic. Prior research has indicated that performance on such tasks can help determine whether a student is likely to understand a particular text based on their knowledge state. The described system is intended for use with middle and high school student populations of native speakers of English. It is designed to handle single reading passages and is not dependent on any corpus or text collection. Here, we describe the system architecture and present an initial evaluation of the system outputs.

Hwang, S., Seo, J., Kim, H., & Lee, G. G. (2026). A multi-agent framework for feature-constrained difficulty control in reading comprehension item generation (arXiv:2605.19316). arXiv. https://arxiv.org/abs/2605.19316

Recent studies in difficulty-controlled reading comprehension item generation have leveraged large language models (LLMs) to produce items by adjusting difficulty-related features. However, existing methods typically rely on a single-agent prompting approach, which often fails to consistently satisfy specified feature constraints, resulting in items that deviate from the target difficulty level. To address this limitation, we introduce MAFIG, a Multiagent Framework for Feature-constrained Item Generation, where multiple LLM agents and feature-specific evaluators collaborate to generate and iteratively revise items based on intended constraints. Furthermore, to verify the efficacy of MAFIG in difficulty control, we propose a method for constructing a sequence of feature constraint sets that yield items with monotonically increasing difficulty. Experimental results demonstrate that MAFIG generates items that adhere to target constraints at a significantly higher rate than baselines, achieving robust difficulty control through the difficulty-calibrated constraint sequence.

Tian, Y., Huynh, L., Christhilf, K., Chakraborty, S., Watanabe, M., Arner, T., & McNamara, D. (2026). Cognitively diverse multiple-choice question generation: A hybrid multi-agent framework with large language models (arXiv:2602.03704). arXiv. https://arxiv.org/abs/2602.03704

Recent advances in large language models (LLMs) have made automated multiple-choice question (MCQ) generation increasingly feasible; however, reliably producing items that satisfy controlled cognitive demands remains a challenge. To address this gap, we introduce Re-QUESTA, a hybrid, multi-agent framework for generating cognitively diverse MCQs that systematically target text-based, inferential, and main idea comprehension. ReQUESTA decomposes MCQ authoring into specialized subtasks and coordinates LLM-powered agents with rule-based components to support planning, controlled generation, iterative evaluation, and post-processing. We evaluated the framework in a large-scale reading comprehension study using academic expository texts, comparing ReQUESTA-generated MCQs with those produced by a single-pass GPT-5 zero-shot baseline. Psychometric analyses of learner responses assessed item difficulty and discrimination, while expert raters evaluated question quality across multiple dimensions, including topic relevance and distractor quality. Results showed that ReQUESTA-generated items were consistently more challenging, more discriminative, and more strongly aligned with overall reading comprehension performance. Expert evaluations further indicated stronger alignment with central concepts and superior distractor linguistic consistency and semantic plausibility, particularly for inferential questions. These findings demonstrate that hybrid, agentic orchestration can systematically improve the reliability and controllability of LLM-based generation, highlighting workflow design as a key lever for structured artifact generation beyond single-pass prompting.

Wang, Y., & Meng, Y. (2026). Optimizing distractor quality in a locally developed second language listening test: Integrating generative AI and psychometric methods. Language Testing, 43(2), 141–164. https://doi.org/10.1177/02655322251400375

This study explores the integration of generative artificial intelligence (GenAI) with human experts to improve the quality of distractors in multiple-choice questions (MCQs) for second language (L2) listening tests. A psychometric analysis of responses from 2267 EFL Chinese undergraduates, using the two-parameter logistic nested logit model (2PLNLM), identified problematic items and distractors. Guided by established distractor design principles, GenAI was applied iteratively to refine these distractors, and GenAI was iteratively used to revise these distractors, with human experts providing ongoing feedback throughout the process. The revised versions were then evaluated by expert judgment and NLP-based cosine similarity analysis. The results indicate that GenAI effectively enhanced distractor quality by maintaining content and structural alignment and ensuring semantic independence. However, it struggled to fully capture listening miscomprehension patterns and contextualized language use. These preliminary findings suggest that GenAI revisions, guided by principle-based prompts and supervised by humans, tend to effectively improve the quality of distractors. This study offers practical insights into the potential and limitations of GenAI in improving L2 listening tests.

Full bibliography

Alsagoafi, A. A., & Alomran, H. S. (2025). Revolutionizing assessment: Leveraging ChatGPT with EFL teachers. World Journal of English Language, 15(6), 385–401.  https://doi.org/10.5430/wjel.v15n6p385

Aryadoust, V., & Wong, J. (2026). How to train your dragon: Evaluating prompting and fine-tuning for GPT-based item generation in L2 listening assessment. Computers and Education: Artificial Intelligence. Advance online publication. https://doi.org/10.1016/j.caeai.2026.100623

Aryadoust, V., Zakaria, A., & Jia, Y. (2024). Investigating the affordances of OpenAI’s large language model in developing listening assessments. Computers and Education: Artificial Intelligence, 6. https://doi.org/10.1016/j.caeai.2024.100204 

Attali, Y., LaFlair, G., & Runge, A. (2023, March 31). A new paradigm for test development [Duolingo webinar series]. Watch webinar

Attali, Y., Runge, A., LaFlair, G.T., Yancey, K., Goodwin, S., Park, Y., & von Davier, A. (2022). The interactive reading task: Transformer-based automatic item generation. Frontiers in Artificial Intelligence, 5, 903077. https://doi.org/10.3389/frai.2022.903077

Belzak, W.C.M., Naismith, B., Burstein, J. (2023). Ensuring fairness of human- and AI-generated test Items. In: N. Wang, G. Rebolledo-Mendez, V. Dimitrova, N. Matsuda, O.C. Santos (Eds.), Artificial Intelligence in education. Communications in Computer and Information Science, 1831. Springer, Cham. https://doi.org/10.1007/978-3-031-36336-8_108 

Bezirhan, U., & von Davier, M. (2023). Automated reading passage generation with Open AI’s large language model. Preprint. https://doi.org/10.48550/arXiv.2304.04616

Bolender, B., Foster, C. & Vispoel, S. (2023). The criticality of implementing principled design when using AI technologies in test development. Language Assessment Quarterly, 20(4-5), 512-519. https://doi.org/10.1080/15434303.2023.2288266 

Bulut, O., & Yildirim-Erbasli, S.N. (2022). Automatic story and item generation for reading comprehension assessments with transformers. International Journal of Assessment Tools in Education, 9, pp.72-87. https://doi.org/10.21449/ijate.1124382

Choi, I. & Zu, J. (2022), The impact of using synthetically generated listening stimuli on test-taker performance: A case study with multiple-choice, single-selection items. ETS Research Report Series, 2022(1), 1–14. https://doi.org/10.1002/ets2.12347

Chun, J. Y. & Barley, N. (2024). A comparative analysis of multiple-choice questions: ChatGPT-generated items vs. human-developed items. In C. A. Chapelle, G. H. Beckett, and J. Ranalli (Eds.), Exploring AI in applied linguistics (pp.118-136). Iowa State University Digital Press. https://bit.ly/TSLL23openbook

Circi, R., Hicks, J., & Sikali, E. (2023). Automatic item generation: Foundations and machine learning-based approaches for assessments. Frontiers in Education, 8:858273.  https://doi.org/10.3389/feduc.2023.858273

Dijkstra, R., Gen¸c, Z., Kayal, S., & Kamps, J. (2022). Reading comprehension quiz generation using generative pre-trained transformers. Pre-print. https://intextbooks.science.uu.nl/workshop2022/files/itb22_p1_full5439.pdf

Fei, Z., Zhang, Q., Gui, T., Liang, D., Wang, S., Wu, W., & Huang, X. (2022). CQG: A simple and effective controlled generation framework for multi-hop question generation. In
Proceedings of the 60th Annual Meeting of the Association for Computational
Linguistics (Volume 1: Long Papers).
Association for Computational Linguistics, 2022. https://aclanthology.org/2022.acl-long.475

Felice, M., Taslimipoor, S., & Buttery, P. (2022). Constructing open cloze tests using generation and discrimination capabilities of transformers. In Findings of the Association for Computational Linguistics: ACL 2022, pp. 1263–1273, Dublin, Ireland. Association for Computational Linguistics. https://arxiv.org/pdf/2204.07237.pdf

Flor, M., Wang, Z., Deane, P., & O’Reilly, T. (2026). Toward an automatic method for generating topical vocabulary test forms for specific reading passages (Research Report No. RR-26-02). ETS. https://doi.org/10.64634/kb8bf328

Ghanem, B., Coleman, L.L., Dexter, J. R., von der Ohe, S. M., & Fyshe, A. (2022). Question generation for reading comprehension assessment by modelling how and what to ask. https://doi.org/10.48550/arXiv.2204.02908

Hwang, S., Seo, J., Kim, H., & Lee, G. G. (2026). A multi-agent framework for feature-constrained difficulty control in reading comprehension item generation (arXiv:2605.19316). arXiv. https://arxiv.org/abs/2605.19316

Kalpakchi, D., & Boye, J. (2021). BERT-based distractor generation for Swedish reading comprehension questions using a small-scale dataset. Paper presented at the 14th International Conference on Natural Language Generation INLG2021. https://arxiv.org/pdf/2108.03973.pdf

Kalpakchi D., & Boye, J. (2023a). Quasi: A synthetic question-answering dataset in Swedish using GPT-3 and zero-shot learning. In T. Alumäe and M. Fishel (Eds.), Proceedings
of the 24th Nordic Conference on Computational Linguistics
(pp.477–491). https://aclanthology.org/2023.nodalida-1.48/

Kalpakchi D., & Boye, J. (2023b). Generation and evaluation of multiple-choice reading
comprehension questions for Swedish.
https://urn.kb.se/resolve?urn=urn:nbn:se:kth:diva-329400

Khademi, A. (2023). Can ChatGPT and Bard generate aligned assessment items? A reliability analysis against human performance. Journal of Applied Learning & Teaching, 6(1), pp.75-80. https://doi.org/10.37074/jalt.2023.6.1.28

Liusie, A., Raina, V., & Gales, M. (2023). “World knowledge” in multiple choice reading
comprehension.
In Proceedings of the Sixth Fact Extraction and VERification Workshop
(FEVER). Association for Computational Linguistics.
https://aclanthology.org/2023.fever-1.5

Ma, W. A., Flor, M., & Wang, Z. (2025). Automatic Generation of Inference-Making Questions for Reading Comprehension. arXiv, June 2025. https://doi.org/10.48550/arXiv.2506.08260

O’Grady, S. (2023). An AI generated test of pragmatic competence and connected speech. Language Teaching Research Quarterly, 37, 188-203. https://doi.org/10.32038/ltrq.2023.37.10

Poon, Y., Wang, Q., Lee, J. S. Y., Lam, Y. Y., & Chu, S. K. W. (2025). PIRLS category-specific question generation for reading comprehension. In Proceedings of the 14th Workshop on Natural Language Processing for Computer Assisted Language Learning (NLP4CALL 2025) (pp. 72–80).  https://hdl.handle.net/10062/107171

Raina, V., & Gales, M. (2022). Multiple-choice question generation: Towards an automated
assessment framework.
https://doi.org/10.48550/arXiv.2209.11830

Raina, V., Liusie, A., & Gales, M. (2023a). Analyzing multiple-choice reading and listening
comprehension tests. https://doi.org/10.48550/arXiv.2307.01076

Raina, V., Liusie, A., & Gales, M. (2023b). Assessing distractors in multiple-choice tests. https://doi.org/10.48550/arXiv.2311.04554

Rathod, A., Tu, T., & Stasaski, K. (2022). Educational multi-question generation for reading
comprehension.
In Proceedings of the 17th Workshop on Innovative Use of NLP for
Building Educational Applications
(pp.216-223). https://aclanthology.org/2022.bea-1.26

Ripoll Y Schmitz, L. M., & Sonnleitner, P. (2025). Evaluating AI-generated vs. human-written reading comprehension passages: An expert SWOT analysis and comparative study for an educational large-scale assessment. Large-scale Assessments in Education, 13(1), Article 20. https://doi.org/10.1186/s40536-025-00255-w

Rodriguez-Torrealba, R., Gracia-Lopez, E., & Garcia-Cabot, A. (2022). End-to-end generation of multiple-choice questions using text-to-text transfer transformer models. Expert Systems With Applications, 208, 118258. https://doi.org/10.1016/j.eswa.2022.118258

Runge, A., Attali, Y., LaFlair, G. T., Park, Y., & Church, J. (2024). A generative AI-driven interactive listening assessment task. Frontiers in Artificial Intelligence, 7, 1474019.
https://doi.org/10.3389/frai.2024.1474019

Sayin, A., & Gierl, M. (2024). Using OpenAI GPT to generate reading comprehension items. Educational Measurement 43(1), 5-18. https://doi.org/10.1111/emip.12590

Shin, I., & Gierl, M. (2022). Generating reading comprehension items using automated processes. International Journal of Testing, 22(3-4), 289-311.
https://doi.org/10.1080/15305058.2022.2070755

Shin, D., & Lee, J. H. (2024). AI-powered automated item generation for language testing. ELT Journal, ccae016. https://doi.org/10.1093/elt/ccae016

Shin, D., Lee, J. H., & Kim, K. (2025). An exploratory study on two automated item generators for generating L2 reading test items. RELC Journal. Advance online publication. https://doi.org/10.1177/00336882251326284

Song, Y., Du, J., & Zheng, Q. (2025). Automatic Item Generation for Educational Assessments: A Systematic Literature Review. Interactive Learning Environments. https://doi.org/10.1080/10494820.2025.2482588

Tan, B., Armoush, N., Mazzullo, E., Bulut, O., & Gierl, M. (2025). A Review of Automatic Item Generation Techniques Leveraging Large Language Models. International Journal of AIG & Testing Education, 12(2), 317–340.  https://doi.org/10.21449/ijate.1602294

Tian, Y., Huynh, L., Christhilf, K., Chakraborty, S., Watanabe, M., Arner, T., & McNamara, D. (2026). Cognitively diverse multiple-choice question generation: A hybrid multi-agent framework with large language models (arXiv:2602.03704). arXiv. https://arxiv.org/abs/2602.03704

von Davier, A. (2023, February 27). Generative AI for test development [a talk given for the Department of Education, University of Oxford].  Watch presentation

Uto, M., Tomikawa, Y., & Suzuki, A. (2023). Difficulty-controllable neural question generation for reading comprehension using item response theory. In Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (pp.119-129). https://aclanthology.org/2023.bea-1.10

Wang, X., Liu, B., & Wu, L. (2023). SkillQG: Learning to generate question for
reading comprehension assessment.
https://doi.org/10.48550/arXiv.2305.04737

Wang, Y., & Meng, Y. (2026). Optimizing distractor quality in a locally developed second language listening test: Integrating generative AI and psychometric methods. Language Testing, 43(2), 141–164. https://doi.org/10.1177/02655322251400375

Yunjiu, L., Wei, W., & Zheng, Y. (2022). Artificial intelligence-generated and human expert-designed vocabulary tests: A comparative study. SAGE Open, 12(1). https://doi.org/10.1177/21582440221082130

Zhang, T., Erlam, R., & de Magalhães, M. (2025). Exploring the dual impact of AI in post-entry language assessment: Potentials and pitfalls. Annual Review of Applied Linguistics, 1–20. https://doi.org/10.1017/S0267190525000030