Past events
Dr Olena Rossi took part in the annual conference of the European Association for Language Testing and Assessment (EALTA) held in June 2026 in Siena, Italy. She gave two presentations and also acted as the discussant for the symposium ‘AI-mediated assessment literacies across stakeholder groups’.
PRESENTATION 1
Beyond perfecting the prompt: Can text-mapping enhance the quality of GenAI-generated B2 reading items? (co-authored with Pep Montcada)
At higher levels of reading proficiency, test-takers are expected to demonstrate the ability to engage with higher-order cognitive processing such as inferencing and evaluation. Previous research using a detailed prompt and advanced item-generation approach found that GenAI was able to replicate the form of higher-order questions but struggled to capture their substance (Rossi & Montcada, 2025). These results indicate that continuously refining the prompt may not be sufficient and that novel approaches to automated item generation may be needed.
This study aimed to investigate whether introducing a text-mapping stage could strengthen construct alignment and improve item quality by addressing two research questions: (1) whether Large Language Models (LLMs) can produce valid B2-level reading text maps – both individually and as an aggregated “panel” – and how these outputs compare to text-mapping by human experts; and (2) whether LLMs generate higher-quality items when guided by a text map. Five LLMs were trained to map two B2 texts, and their mappings were compared with those of human experts. Using these text maps, multiple-choice items were generated, which then underwent expert review and trialling with B2 learners. The presentation reported findings for the text-mapping stage of the project.
Download presentation slides
PRESENTATION 2
Conceptualising AI literacy for the language assessment profession (co-authored with Stefan O’Grady and Nazlinur Gokturk)
The rapid rise in the use of Artificial Intelligence (AI) for language assessment has amplified the need for practitioners to develop AI literacy. While current conceptualisations of language assessment literacy (LAL) acknowledge digital/technological literacy, the field still lacks a theoretically grounded conceptualisation of AI literacy specific to language assessment. This paper proposes a definition of AI literacy tailored to assessment practitioners and positions it as an essential component of LAL in contemporary large-scale and classroom assessment contexts.
The paper draws on two widely accepted definitions of AI literacy – by Long and Magerko (2020) and Chiu (2025) – and integrates these with UNESCO’s (2024) AI Competency Framework for Teachers, particularly its developmental AI competency progression of Acquire–Deepen–Create. Building on these foundations and on recent discussions of AI uses, opportunities, risks, and ethical considerations in language assessment, the paper outlines what AI literacy entails for two major stakeholder groups: (1) practitioners working in operational testing (item development, automated delivery, scoring, proctoring, analytics, validation), and (2) teachers conducting classroom-based assessment (task creation, feedback, analytics, adaptive systems, chatbots). It then proposes progression levels describing how AI literacy develops across professional stages in the two contexts.
This paper’s contribution is threefold: (a) providing the first field-specific definition of AI literacy for language assessment, (b) articulating differentiated literacy demands across stakeholder groups, and (c) conceptualising AI literacy as an integral dimension of modern LAL. The paper aims to stimulate debate on how assessment professionals can be prepared for responsible and theoretically informed AI use.
Download presentation slides
On 26th May 2026, Dr Olena Rossi gave an invited talk at the webinar organised by the Asian Association for Language Assessment (AALA).
Title: Generative AI for item writing: outstanding questions and future directions
The presentation examined key challenges in the use of generative AI for language test item writing as well as future directions for AI use. It addressed two unresolved and seldom discussed issues that impact on the quality of AI-produced items: the lack of high-quality training data for fine-tuning large language models, and the absence of valid automated item evaluation metrics capable of replacing human judgement. It then explored ways to enhance AI performance, arguing that improving output quality will increasingly depend on redesigning item-generation workflows rather than on refining prompts alone.
Download presentation slides
In January 2026, Dr Olena Rossi represented the European Association for Language Testing and Assessment (EALTA) at the AI-Forward Winter School on Artificial Intelligence for Language Learning, Teaching and Assessment, organised by the University for Foreigners of Perugia in cooperation with ALTE (Association for Language Testers in Europe). Olena Rossi gave a keynote lecture titled “AI Advancements in Language Assessment: Opportunities and Risks”.
The lecture provided a structured overview of recent advances in AI and its applications in foreign language assessment, exploring key areas including automated item generation, computer-adaptive testing, spoken dialogue systems for speaking test delivery, remote proctoring, automated scoring, automated feedback, and predictive analytics. Alongside the opportunities associated with these applications, the lecture highlighted the limitations and risks of AI use, such as bias, lack of transparency, data privacy concerns, and over-reliance on automated systems.
Download presentation slides
Dr Olena Rossi took part in the annual conference of the European Association for Language Testing and Assessment (EALTA) held in May 2025 in Salzburg, Austria. She co-authored two presentations:
PRESENTATION 1
Generating reading test items with AI: Targeting higher-order cognitive processing (with Pep Montcada)
Presentation abstract
While generative AI (GenAI) can produce text-based items of various types, the items often target information recall rather than higher-order thinking skills like inferencing, evaluation, or synthesis (Voss, 2024; Xi, 2024). This study, a collaboration between an assessment researcher and a large-scale test provider, explores the potential of GenAI for producing multiple-choice (MC) items targeting higher-order thinking skills for the reading comprehension test of the English B2 certification exam of the Official Schools of Languages in Catalonia.
The test uses authentic texts, and six retired texts were selected for the study. A detailed prompt based on the test specifications was developed for item generation using ChatGPT-4o. The progressive prompting technique was employed to refine outputs iteratively. From 30 items generated per text, 5-10 items were shortlisted for independent evaluation by two professional reviewers. In a subsequent reviewer meeting, two best tasks were selected for trialling. These tasks were administered to 775 students from 30 Catalonian schools under exam-like conditions. The students also gave feedback on item clarity and difficulty.
Test scores were analysed using Classical Test Theory (CTT) and Item Response Theory (IRT). Item evaluations by the reviewers and students were analysed qualitatively. Results showed that ChatGPT4o was able to generate a number of usable items targeting higher-order thinking skills, but human involvement was critical for item generation, selection and evaluation.
The presentation will discuss the implications for using GenAI in test development, including recommendations for prompt design, troubleshooting strategies, and the required level of human intervention.
Download presentation slides
The study has since been published as a pre-print:
Rossi, O, & Montcada, J. M. (2025). Generating multiple-choice items for a B2 English reading test with GPT-4: targeting higher-order cognitive processing. Preprint. https://doi.org/10.35542/osf.io/qjfpb_v1
PRESENTATION 2
AI in Language Testing: Trends, Gaps and Future Directions (with Stefan O’Grady and Nazlinur Gokturk)
Presentation abstract
Artificial Intelligence (AI) is transforming language assessment through applications such as automated scoring, item generation, test administration, and more. With research on AI in language assessment increasing exponentially, language assessment specialists have been left “scrambling” to remain up to date (Xi, 2023, p. 357). To that end, obtaining a comprehensive overview of this work is essential. This presentation reports on a scoping review of empirical studies published between 2019 and 2024, addressing the research question: What is known from the existing literature about the use of AI in language testing and assessment?
The present study followed PRISMA reporting standards for review and meta-analyses (Page, 2021). To identify relevant studies, we searched four electronic databases (SCOPUS, Web of Science, ERIC, and LLBA) using carefully developed search terms. The initial search yielded 4,055 results, which were narrowed down through duplicate removal and eligibility coding to a final set of 312 relevant studies. These studies were analysed and summarised using a charting form adapted from Arksey and O’Malley (2004) to facilitate comparisons. Patterns and trends in the data were then examined to provide a comprehensive overview.
The presentation will discuss geographical distribution, key themes, methodologies, and findings from this five-year body of research, offering invaluable insights into how AI tools shape language assessment contexts. Additionally, the presentation will identify gaps in the current knowledge base and propose directions for future research. The presentation will be of interest to researchers and practitioners with an interest in the various applications of AI in language assessment.
The study has since been published in the journal Research Synthesis in Applied Linguistics:
O’Grady, S., Gokturk, N. & Rossi, O. (2026): Artificial Intelligence in language assessment research: A Scoping review. Research Synthesis in Applied Linguistics, advance online publication. https://doi.org/10.1080/29984475.2026.2633389 (free access)
In March 2025, Dr Olena Rossi, in collaboration with Pep Montcada, presented their recent study on using ChatGPT for True/False reading comprehension item generation. The presentation provided practical recommendations on producing True/False items with Generative AI.
Presentation abstract:
Generative AI (GenAI) is increasingly being used to develop language test items across various formats. True/False (TF) items are popular for assessing reading comprehension in classroom contexts. This presentation shares the findings of a study that examined the potential of GenAI for generating TF items for B1-B2 reading comprehension tests. The presentation will offer practical guidance on using GenAI to create high-quality TF items.
The study focussed on reading comprehension tests for the English B1 and B2 certification exams administered by the Official Schools of Languages in Catalonia. These exams use authentic texts, and for this study, twelve retired texts (six per proficiency level) were selected. A detailed prompt based on the test specifications was developed for item generation using ChatGPT-4o. A progressive prompting technique was applied to iteratively refine the outputs. From approximately 30 items generated per text, 5-10 were shortlisted for independent evaluation by two professional reviewers. In a subsequent reviewer meeting, two best tasks per proficiency level (four in total) were selected for trialling.
These tasks were then administered to 811 (B1) and 775 (B2) students from 30 Catalonian schools under exam-like conditions. The students also gave feedback on item clarity and difficulty. Test scores were analysed using Classical Test Theory (CTT) and Item Response Theory (IRT). Item evaluations by the reviewers and students were analysed qualitatively.
This presentation will discuss the study’s findings and their implications for using GenAI in TF item development. Key takeaways will include recommendations for prompt design, troubleshooting strategies, and the extent of human intervention required to ensure high-quality item generation.
Download presentation slides
Watch presentation
In December 2024, Dr Olena Rossi presented at the BAAL TEASIG Webinar on Item Generation and AI, alongside Dr Alina von Davier and Dr Vahid Aryadoust.
Olena Rossi presented on AUTOMATED ITEM GENERATION: AN ITEM WRITER PERSPECTIVE.
Presentation abstract:
Generative AI is transforming language test development, prompting questions about the future of item writers and their evolving role. This presentation focusses on item writers, examining their perceptions, current practices, and professional trajectory. Drawing on recent studies, it explores item writer familiarity with generative AI tools and their attitudes toward integrating these technologies into their work.
The presentation highlights the dual potential of generative AI in item writing: as a productivity booster and as a collaborator in innovative test design. It emphasizes a “human-in-the-lead” approach, where AI acts as a “co-pilot” to support creativity and efficiency of human item writers. Through a discussion of case studies, the presentation demonstrates how AI can be effectively used to generate high-quality items while maintaining human oversight. The presentation concludes with a vision of item writers as technologically adept professionals equipped with prompt engineering and multimodal design skills, ensuring the production of high-quality items in an AI-enhanced landscape.
Download presentation slides
Dr Olena Rossi gave a workshop on using ChatGPT to generate tasks for EAP (English for Academic Purposes) reading and listening assessments. The workshop was held as part of the BALEAP Assessment Roadshows. The workshop was delivered online on 25 July 2024.
In this hands-on workshop, participants learned practical strategies for leveraging ChatGPT in creating reading and listening tasks for EAP assessments. They delved into ChatGPT’s capabilities for generating texts for reading assessments tailored to various genres and proficiency levels. They also explored how ChatGPT can generate scripts for listening tasks, followed by the generation of audio using human-like AI voices.
The workshop also explored ChatGPT’s ability to generate different types of assessment items. We investigated whether ChatGPT can cater to different levels of cognitive processing, ranging from basic comprehension tasks like word recognition and sentence-level comprehension, to higher-order processes such as inferencing, text-level representation and intertextual understanding.
Watch Part 1 of the workshop
Watch Part 2 of the workshop
Download workshop slides
Dr Olena Rossi convened an inaugural meeting of the EALTA special interest group – Artificial Intelligence for Language Assessment. The SIG meeting was held as part of the EALTA Annual Conference. 6 June 2024, Belfast, UK.
The aim of the AI in LA SIG is to provide a dedicated forum for discussing and sharing knowledge and practices in the use of AI in language testing and assessment. This SIG has its own website where the members and convenors started a digital library and media library.
During the meeting, Olena Rossi gave a presentation ‘Item writing with generative AI: Current issues and future directions’.
Abstract:
While there is widespread acknowledgment of generative AI’s capacity to create language test items, numerous unresolved questions persist. Discussions frequently centre on AI’s tendency to generate false information, the presence of biases in AI-generated text, and the challenge of crafting items that effectively evaluate higher-order thinking skills. In this presentation, however, I aim to explore three different challenges confronting automated item generation today that I see as fundamental: 1) the lack of high-quality training data; 2) the absence of robust automatic evaluation metrics for appraising item quality, and 3) insufficient interdisciplinary collaboration between natural language processing experts and language testing professionals. Moreover, I will offer insights into the potential future trajectory of generating language test items with AI.
Dr Olena Rossi gave a plenary talk at the 2024 International Conference ‘Language Education 4.0: A Paradigm Shift towards Action-Oriented Approach, Artificial Intelligence Integration and Beyond’, 31 May 2024, Bilkent University, Ankara, Türkiye
Olena Rossi’s talk was entitled: “Assessment of language through AI: Opportunities, challenges, and future directions”
Abstract:
In this talk, I will focus on using artificial intelligence for language assessment. The talk will start with an overview of opportunities and challenges of applying AI to various areas of language assessment. These include automated scoring, automated feedback, remote proctoring, computer-mediated interactive testing of productive skills, as well as text and item generation. The second part of the talk will cast a deeper look into the prospects of using generative AI tools such as ChatGPT, Bing as well as commercial AI-powered item writing platforms for producing language test items. In discussing this topic, I will summarise recent developments in the field, talk about my personal experience producing test items with AI, and touch upon some recent research. I will dwell both on the great opportunities AI has created for the classroom as well as larger-scale testing, and on the substantial challenges the use of AI has presented.
Dr Olena Rossi gave a presentation at the IATEFL TEASIG Online Event ‘Developing assessment tasks for the classroom’, September 2023
Olena Rossi’s talk was entitled “Using technology to write language test items: Opportunities and challenges”
Abstract:
In this presentation, I will focus on how technology can be used in the process of producing language test items. I will first discuss the role of corpus tools in item development. This includes using online profilers such as EDI Papyrus or Lextutor to grade test materials, as well as using L1 and learner corpora to select salient language for testing, to write naturally-sounding items, and to create strong distractors. I will then proceed to discuss the prospects of using AI tools such as ChatGPT and Bing for text and item generation. I will dwell both on the great opportunities Artificial Intelligence has created for the classroom as well as larger-scale testing, and on the substantial challenges the use of AI has presented.

