• A
  • A
  • A
  • АБВ
  • АБВ
  • АБВ
  • A
  • A
  • A
  • A
  • A
Обычная версия сайта
  • RU
  • EN
  • HSE University
  • Publications
  • Articles
  • Large Language Model-Based Automated Item Generation in STEM Assessments: Historical Mapping and a Scoping Review of Empirical Studies
  • RU
  • EN
Расширенный поиск
Высшая школа экономики
Национальный исследовательский университет
Priority areas
  • business informatics
  • economics
  • engineering science
  • humanitarian
  • IT and mathematics
  • law
  • management
  • mathematics
  • sociology
  • state and public administration
by year
  • 2028
  • 2027
  • 2026
  • 2025
  • 2024
  • 2023
  • 2022
  • 2021
  • 2020
  • 2019
  • 2018
  • 2017
  • 2016
  • 2015
  • 2014
  • 2013
  • 2012
  • 2011
  • 2010
  • 2009
  • 2008
  • 2007
  • 2006
  • 2005
  • 2004
  • 2003
  • 2002
  • 2001
  • 2000
  • 1999
  • 1998
  • 1997
  • 1996
  • 1995
  • 1994
  • 1993
  • 1992
  • 1991
  • 1990
  • 1989
  • 1988
  • 1987
  • 1986
  • 1985
  • 1984
  • 1983
  • 1982
  • 1981
  • 1980
  • 1979
  • 1978
  • 1977
  • 1976
  • 1975
  • 1974
  • 1973
  • 1972
  • 1971
  • 1970
  • 1969
  • 1968
  • 1967
  • 1966
  • 1965
  • 1964
  • 1963
  • 1958
  • More
Subject
News
September 4, 2026
Time to Showcase Your Research: Applications Are Now Open for Student Research Paper Competition 2026
Taking part in the Student Research Paper Competition (SRPC) gives you an opportunity to present your research to experts, receive an independent assessment, and determine the future direction of your work. The competition is open to students graduating in 2026 not only from HSE University but from universities in Russia and abroad. Papers may be submitted in Russian and English, and in some fields also in French, German, and Spanish.
September 4, 2026
‘Hedgehog Versus ‘Relatives: Researchers Measure How the Brain Responds to Unexpected Words During Natural Speech
Russian neurophysiologists, including researchers from HSE University, have demonstrated the feasibility of using event-related fields (ERFs) to study brain activity during natural speech perception. The researchers showed that this approach can be applied not only to individual words but also to continuous speech. Their findings indicate that words whose meanings differ significantly from the preceding context require longer processing times. The study also reveals that the brain processes function words in two stages: first, it identifies their grammatical role and then uses this information to predict the next word. The study has been published in Frontiers in Human Neuroscience.
August 25, 2026
Scientists Develop Algorithm for More Reliable Processors in Data Centres
Researchers from HSE MIEM and Samara University have developed the LRF-3D algorithm to automatically bypass idle nodes in three-dimensional networks-on-chip. Thanks to its hierarchical architecture, the algorithm outperforms existing solutions in both speed and path accuracy, improving processor reliability for use in data centres, supercomputers, and AI computing. The source code and test results are publicly available.

 

Have you spotted a typo?
Highlight it, click Ctrl+Enter and send us a message. Thank you for your help!

Publications
  • Books
  • Articles
  • Chapters of books
  • Working papers
  • Report a publication
  • Research at HSE

?

Large Language Model-Based Automated Item Generation in STEM Assessments: Historical Mapping and a Scoping Review of Empirical Studies

JOURNAL OF EDUCATIONAL TECHNOLOGY DEVELOPMENT AND EXCHANGE. 2026. Vol. 19. No. 2. P. 141–165.
Omopekunola M.

Educational assessments, from low-stakes classroom tests to high-stakes national examinations, require item pools that are valid, fair, and secure. Automated Item Generation (AIG) aims to efficiently produce large pools of calibrated test items. This paper adopts a two-part design: (1) a brief historical mapping situating LLM-based AIG within the broader AIG trajectory; and (2) a scoping review of empirical studies on LLM-based AIG for STEM assessments, published between January 2022 and January 2026. A structured search of ERIC, Lens and OpenAlex yielded 1,267 records; after deduplication and screening, 7 studies were retained for synthesis. In all studies, LLMs were primarily used to draft stems, keys, distractors, and explanations by instruction-tuned prompting, sometimes enhanced with retrieval and human-in-the-loop review. Empirical evidence on item quality is generally promising. Multiple investigations have documented acceptable expert evaluations and, in a subset of studies, psychometric properties comparable to those of human-authored items. Nevertheless, recurrent limitations have been observed, including factual inaccuracies, construct drift, low calibration of item difficulty, and variable distractor plausibility. Few studies reported robust fairness audits or provided reproducible details, such as complete prompts and decoding settings. In general, LLM-based AIG can substantially increase throughput in STEM item development, but high-stakes deployment requires layered validation protocols (expert review, pilot testing, psychometrics, and bias audits) and governance controls to ensure traceability and item security.

Research target: Education Psychology
Language: English
DOI
Text on another site
Keywords: automated item generationLarge language models (LLM)prompt engineeringHigh-stakes assessmentPsychometric validation
Similar publications
Self-esteem and self-reported personality traits: When attitudes toward traits make a difference
Shchebetenko S., De Marchis G., / Series PsyArXiv "PsyArXiv". 2026.
Personality questionnaires do not directly record recurring conduct. Rather, they elicit task-bounded reports of personality attributions: how respondents characterise themselves using personality-descriptive content. From this perspective, established associations between self-esteem and self-reported personality may depend on the evaluative meaning attached to that content. We tested whether attitudes toward the Big Five traits moderated associations between self-esteem and corresponding personality reports ...
Added: September 4, 2026
Lower resting-state EEG mean eigenvector centrality is associated with higher math anxiety
Pavlova A., Малых С. Б., Адамович Т. В., Frontiers in Human Neuroscience 2026 No. 20 Article 1843779
This study aimed to investigate resting-state neural correlates of math anxiety in high school students using a graph-theoretical approach. Sixty 10th-grade students (35 females; mean age = 16.15 years) participated in the study. Resting-state EEG data were acquired using a 32-channel wireless dry-sensor system (Cognionics Quick-32 series). Math anxiety was assessed using the Abbreviated Math Anxiety Scale. Mean ...
Added: September 3, 2026
Логика и парадоксы управления современным образованием: Всероссийский круглый стол (Москва, МГУ имени М. В. Ломоносова, социологический факультет, 8 июня 2026 г.)
М.: МАКС Пресс, 2026.
The book consists of materials submitted by Russian and Foreign scientists to the Organizing Committee of the All-Russian Round Table “Logic and Paradoxes of Modern Education Manage ment”. The round table was held at the Faculty of Sociology of Lomonosov Moscow State University on June 08, 2026. The texts are published in author's edition. Organizing Committee`s point of ...
Added: September 2, 2026
Informal Workers’ Capabilities and Recognition of Prior Learning. Current Practice and Alternatives
Mehrotra S. K., Sharma A., Mehrotra V. S., Economic and Political Weekly 2025 Vol. 60 No. 51
Recognition of Prior Learning, a key component of the Skill India Mission under the National Policy for Skill Development and Entrepreneurship, 2015, seeks to formalise the skills of India’s vast informal workforce, which accounts for over 91% of the total employment. While RPL can enhance employability, economic mobility, and workforce productivity, its implementation remains constrained ...
Added: September 1, 2026
Beyond the Binary: A Theoretical Exploration of Gendered Experiences for Women in Nigerian STEM Education
Asia and Africa today 2026 Vol. 6 P. 80–86
This theoretical article examines the persistent problem of the “leaky pipeline” for women in Nigerian higher education, where increased representation in male-dominated STEM fields is undermined by enduring stereotypes and socio-cultural barriers. By employing a critical theoretical analysis and interpretive synthesis, it aims to demonstrate how the concept of reflexive modernization explains why stereotype threats ...
Added: August 31, 2026
Обществознание
Ермоленко Г. А., Кожевников С. Б., М.: Эксмо, 2026.
Socio‑humanitarian knowledge helps to master the social norms and values of Russian civilization. Each module of social studies is linked to a specific sphere of human social life, to activities aimed at fulfilling a person’s spiritual needs and social interests. Recognizing the social nature of one’s own needs is an important step towards understanding oneself, ...
Added: August 31, 2026
The Use of Outsourcing in Schools under Different Political Environments
Panova A., Ostrovnaya M., Demokratizatsiya: The Journal of Post-Soviet Democratization 2026 Vol. 34 No. 2 P. 199–226
Local authorities can crucially affect the economic performance of public organizations. Using data on the Moscow region, we analyze the unique case of a huge, socio-economically homogeneous region with differing political environments. We examine how the political characteristics of local authorities influence state schools’ decisions on whether to outsource non-core activities, in this case food ...
Added: August 27, 2026
Демократия в эпоху социальных трансформаций: два различных подхода
Medushevsky A. N., Liberal.ru 2026
The Liberal Mission Foundation continues to introduce readers to new concepts of democracy and relevant publications. This discussion focuses on the challenges facing democracy in the era of globalization, as addressed in two recent books: *The Backsliders: Why Leaders Undermine Their Own Democracies* by S.C. Stokes (2025) and *Critical Theories and the Postsocialist Challenge: Race, ...
Added: August 26, 2026
Искусственный интеллект и этика будущего
Medushevsky A. N., Liberal.ru 2026
Today, more than ever before, the contribution of new technologies shapes the conditions of humanity's existence as a biological species, prompting a debate about its prospects for survival. This process of transformation—which spans every sphere of social and moral regulation, from family relations and child-rearing to the very meaning of life—is dividing public opinion between ...
Added: August 26, 2026
Добродетель и порок: этические кодексы маргинальных сообществ современной России. Часть 2
Medushevsky A. N., Вопросы теоретической экономики 2026 № 2 С. 146–164
It is generally accepted that the basis of social and legal stability in any society is a consensus on vir tue — the fundamental values of proper behavior, ways of maintaining and reproducing them. And, conversely, the destruction of this consensus is a sign of the loss of stability and divergence of positions of social ...
Added: August 26, 2026
Development and validation of the Mental Health Self-Care Scale for the general adult population
Mikhaylova O., Zhyrgalbek J., Sofia Yanis et al., Public Health 2026 Vol. 258 Article 106405
Objectives: To develop and psychometrically validate the Mental Health Self-Care Scale (MHSCS) for assessing mental health self-care behaviours in the general adult population. Study design: Cross-sectional survey. Methods: A 62-item initial pool grounded in the updated Middle Range Theory of Self-Care was refined through expert review (n = 11) and cognitive interviewing (n = 24), then administered to 600 Russian ...
Added: June 25, 2026
Use Case 5: LLM-driven creation of natural hazard geodatabase from digital mass media
Derkacheva A., Sakirkina M., Kraev G. et al., , in: AI for good innovate for impact report 2025.: Geneva: International Telecommunication Union, 2025. P. 167–169.
Added: May 26, 2026
Об идеологических предвзятостях генеративного ИИ: Российско-украинский конфликт в репрезентации ChatGPT
Baysha O., Trofimov V., Российская школа связей с общественностью 2026 № 40 С. 171–191
A growing number of scholars are warning about the dangers of the reproduction by generative AI of socio-political and ideological biases absorbed by models from the texts on which they were trained. If a given model was trained on Western media texts, it may generate narratives that reproduce West centric views of world events. This ...
Added: April 21, 2026
Сопоставление номенклатур товаров ресторанов и поставщиков с помощью LLM — Case Study для ресторанного холдинга
Jin S., Panfilov P., Сулейкин А. С., Труды Института системного программирования РАН 2025 Т. 37 № 6 С. 163–176
In the modern restaurant business, accurate mapping of product nomenclatures between restaurants and suppliers is a critical task. Effective inventory management and procurement optimization directly impact business profitability. With the increase in suppliers and product variety, traditional mapping methods become less efficient. This study proposes using large language models (LLM) to automate and improve the ...
Added: April 17, 2026
Learning When to Personalize: LLM Based Playlist Generation via Query Taxonomy and Classification
Buzaev F., Пугачёва Д. В., Sukharev I. et al., Transactions of the Association for Computational Linguistics 2026 P. 51–57
Playlist generation based on textual queries using large language models (LLMs) is becoming an important interaction paradigm for music streaming platforms. User queries span a wide spectrum from highly personalized intent to essentially catalog-style requests. Existing systems typically rely on non-personalized retrieval/ranking or apply a fixed level of preference conditioning to every query, which can ...
Added: April 7, 2026
Large Language Models as Political Actors: Cultural Bias and Epistemic Power
Seredkina E., Seletkova G., Mikhailovsky A., Technology and Language 2026 Vol. 7 No. 1 P. 63–79
The rapid diffusion of Large Language Models (LLMs) into socially and politically sensitive domains raises critical questions about the nature and origins of political bias in artificial intelligence. While existing research often treats bias as a technical flaw to be minimized, this article advances a broader philosophical and cultural interpretation of LLM bias as an ...
Added: April 1, 2026
Validating the Russian Adult Prosocialness Behavior Scale: Weighted Factor Analysis, Sex Invariance, and Normative Benchmarks
Mikhaylova Oxana, Bochaver Alexandra, Current Psychology 2026 Vol. 45 Article 712
Prosocial behavior measures validated across diverse cultural contexts remain limited. We validated the Adult Prosocialness Behavior Scale (APBS) in 7,965 Russian adults (52.0% women; Mage = 42.4, SD = 12.1) using sampling weights to approximate population representativeness. Split-sample analyses supported a two-factor structure with correlated Prosocial Actions and Prosocial Feelings dimensions (χ² = 2,754.87, CFI = .978, TLI ...
Added: March 11, 2026
Промпт-инжиниринг как ключевая компетенция в образовании: сущность, особенности и подходы к оцениванию
Davlatova M., Сперанская М. В., Высшее образование в России 2026 Т. 35 № 2 С. 53–73
In the context of the rapid development of Generative Artificial Intelligence (GenAI), prompt engineering is becoming a key competence for effective interaction with large language models in educational settings. However, the lack of a unified understanding of its nature, structure, and assessment tools complicates its integration into educational practice.    The aim of this study is to ...
Added: March 7, 2026
Can Large Language Models Develop High-Stakes Physics Exam Items? A Comprehensive Study of Cognitive and Psychometric Efficacy
Moses Oluoke Omopekunola, Elena Yu. Kardanova, Journal of Science Education and Technology 2026 Vol. 35 P. 933–943
High-stakes assessment is crucial for evaluating student performance and making significant educational decisions. Traditionally, the development of test items for such examinations has relied on manual development by subject matter experts. However, Automated Item Generation (AIG) using Large Language Models (LLMs) has emerged as a promising alternative, though systematic research on their application in high-stakes ...
Added: January 16, 2026
Многоаспектная оценка методов адаптации токенизатора для больших языковых моделей на русском языке
Андрющенко Г. Д., Godunova M., Иванов В. В. et al., Доклады Российской академии наук. Математика, информатика, процессы управления (ранее - Доклады Академии Наук. Математика) 2025 Т. 527 С. 320–331
Large language models (LLMs) pretrained on English-centered corpora have biases and perform sub-optimally on other natural languages. Adaptation of LLMs vocabulary provides a resource-efficient way to improve the quality of a pretrained model. Previously proposed adaptation techniques focus on performance (accuracy) and size metrics (fertility), ignoring other aspects in comparison, such as inference latency, compute ...
Added: January 15, 2026
Generating and Debugging Java Code using LLMs based on Associative Recurrent Memory
Vasilevsky V., Alexandrov D., Proceedings of the Institute for System Programming of the RAS 2025 Vol. 37 No. 5 P. 173–182
Automatic code generation by large language models (LLMs) has achieved significant success, yet it still faces challenges when dealing with complex and large codebases, especially in languages like Java. The limitations of LLM context windows and the complexity of debugging generated code are key obstacles. This paper presents an approach aimed at improving Java code generation and debugging. ...
Added: December 26, 2025
Detecting Ethnic Conflict in Social Media with Transformers and Augmented Data
Koltsova O., Surkov A., Procedia Computer Science 2025 Vol. 258 P. 2382–2390
Chest X-ray pathology prediction play a very important role in early disease detection, enabling timely intervention and improving patient outcomes. Detection of ethnic conflict mentioning, discussion, or verbal participation therein in user-generated content is a socially important task, as such content has been proven related to ethnic clashes on the ground. Yet this task has not been ...
Added: November 28, 2025
SIGNAL: Dataset for Semantic and Inferred Grammar Neurological Analysis of Language
Komissarenko A., Voloshina E., Чевелева А. Н. et al., Scientific data 2025 Vol. 12 No. 1 Article 1687
Recently, the idea of brain-model alignment has been the topic of several influential works. However, most of previous studies were based on datasets collected during regular reading tasks where the subjects were not exposed to processing linguistic incongruencies, and stimuli were not controlled for key linguistic properties. Meanwhile, interpretability studies of Large Language Models pay ...
Added: November 18, 2025
Исследования благополучия с помощью передовых методов обработки естественного языка (NLP): перспективы и ограничения
Voevodina E., Современная зарубежная психология 2025 Т. 14 № 3 С. 172–181
Context and relevance. Well-being research faces methodological limitations of conventional psychometric measures, criticized for poor ecological validity, limited information yield, and inadequate capture of multidimensional construct of well-being. Advanced natural language processing (NLP) technologies offer solutions to these constraints. Objective. To evaluate opportunities and challenges of transformer-based NLP for well-being research. Methods and materials. We conducted an analytical review of ...
Added: October 9, 2025
  • About
  • About
  • Key Figures & Facts
  • Sustainability at HSE University
  • Faculties & Departments
  • International Partnerships
  • Faculty & Staff
  • HSE Buildings
  • HSE University for Persons with Disabilities
  • Public Enquiries
  • Studies
  • Admissions
  • Programme Catalogue
  • Undergraduate
  • Graduate
  • Exchange Programmes
  • Summer University
  • Summer Schools
  • Semester in Moscow
  • Business Internship
  • Research
  • International Laboratories
  • Research Centres
  • Research Projects
  • Monitoring Studies
  • Conferences & Seminars
  • Academic Jobs
  • Yasin (April) International Academic Conference on Economic and Social Development
  • Media & Resources
  • Publications by staff
  • HSE Journals
  • Publishing House
  • iq.hse.ru: commentary by HSE experts
  • Library
  • Economic & Social Data Archive
  • Video
  • HSE Repository of Socio-Economic Information
  • HSE1993–2026
  • Contacts
  • Copyright
  • Privacy Policy
  • Site Map
Edit