• A
  • A
  • A
  • АБВ
  • АБВ
  • АБВ
  • A
  • A
  • A
  • A
  • A
Обычная версия сайта
  • RU
  • EN
  • HSE University
  • Publications
  • Articles
  • Wrong Answers Only: Distractor Generation for Russian Reading Comprehension Questions Using a Translated Dataset
  • RU
  • EN
Расширенный поиск
Высшая школа экономики
Национальный исследовательский университет
Priority areas
  • business informatics
  • economics
  • engineering science
  • humanitarian
  • IT and mathematics
  • law
  • management
  • mathematics
  • sociology
  • state and public administration
by year
  • 2027
  • 2026
  • 2025
  • 2024
  • 2023
  • 2022
  • 2021
  • 2020
  • 2019
  • 2018
  • 2017
  • 2016
  • 2015
  • 2014
  • 2013
  • 2012
  • 2011
  • 2010
  • 2009
  • 2008
  • 2007
  • 2006
  • 2005
  • 2004
  • 2003
  • 2002
  • 2001
  • 2000
  • 1999
  • 1998
  • 1997
  • 1996
  • 1995
  • 1994
  • 1993
  • 1992
  • 1991
  • 1990
  • 1989
  • 1988
  • 1987
  • 1986
  • 1985
  • 1984
  • 1983
  • 1982
  • 1981
  • 1980
  • 1979
  • 1978
  • 1977
  • 1976
  • 1975
  • 1974
  • 1973
  • 1972
  • 1971
  • 1970
  • 1969
  • 1968
  • 1967
  • 1966
  • 1965
  • 1964
  • 1963
  • 1958
  • More
Subject
News
October 8, 2026
HSE Experts Take Part in 23rd Annual Meeting of Valdai Discussion Club
The 23rd Annual Meeting of the Valdai Discussion Club was held from September 28 to October 1, 2026 under the theme ‘Responsibility for the Future: Limits of the Possible, or Limitless Possibilities?’ The forum brought together 120 experts from 40 countries, including representatives of China, the United States, India, Brazil, the United Kingdom, Germany, Egypt, Iran, and Japan.
October 7, 2026
‘Our Team Consists of True Leaders in Their Respective Academic Disciplines
The HSE International Centre of Decision Choice and Analysis studies a wide range of methods for analysing decision-making and possible scenarios for the development of natural, socio-economic, and political phenomena using various mathematical models. The application of advanced mathematical methods to forecasting helps to prevent negative outcomes and avoid erroneous decisions. The HSE News Service spoke to the centre’s director, Prof. Fuad Aleskerov, about its work.
October 6, 2026
International N5 Symposium ‘Neural Networks and Nonlinearity in Nizhny Novgorod Brings Together Scientists from Russia and Serbia
The International N5 Symposium ‘Neural Networks and Nonlinearity in Nizhny Novgorod’ was held at the Nizhny Novgorod House of Scientists from September 23 to 26. The event was organised by HSE University–Nizhny Novgorod and the Nizhny Novgorod House of Scientists, with the participation of Sberbank and the Institute of Physics Belgrade. The symposium was held for the second time: the first conference took place in 2025 and attracted considerable interest from the academic community.

 

Have you spotted a typo?
Highlight it, click Ctrl+Enter and send us a message. Thank you for your help!

Publications
  • Books
  • Articles
  • Chapters of books
  • Working papers
  • Report a publication
  • Research at HSE

?

Wrong Answers Only: Distractor Generation for Russian Reading Comprehension Questions Using a Translated Dataset

Journal of Language and Education. 2024. Vol. 10. No. 4. P. 56–70.
Login N.

Background: Reading comprehension questions play an important role in language learning. Multiple-choice questions are a convenient form of reading comprehension assessment as they can be easily graded automatically. The availability of large reading comprehension datasets makes it possible to also automatically produce these items, reducing the cost of development of test question banks, by fine-tuning language models on them. While English reading comprehension datasets are common, this is not true for other languages, including Russian. A subtask of distractor generation poses a difficulty, as it requires producing multiple incorrect items.

Purpose: The purpose of this work is to develop an efficient distractor generation solution for Russian exam-style reading comprehension questions and to discover whether a translated English-language distractor dataset can offer a possibility for such solution.

Method: In this paper we fine-tuned two pre-trained Russian large language models, RuT5 and RuGPT3 (Zmitrovich et al, 2024), on distractor generation task for two classes of summarizing questions retrieved from a large multiple-choice question dataset, that was automatically translated from English to Russian. The first class consisted of questions on selection of the best title for the given passage, while the second class included questions on true/false statement selection. The models were assessed automatically on test and development subsets, and true statement distractor models were additionally evaluated on an independent set of questions from Russian state exam USE.

Results: It was observed that the models surpassed the non-fine-tuned baseline, the performance of RuT5 model was better than that of RuGPT3, and that the models handled true statement selection questions much better than title questions. On USE data models fine-tuned on translated dataset have shown better quality than that trained on existing Russian distractor dataset, with T5-based model also beating the baseline established by output of an existing English distractor generation model translated into Russian.

Conclusion: The obtained results show the possibility of a translated dataset to be used in distractor generation and the importance of the domain (language examination) and question type match in the input data.

Research target: Philology and Linguistics
Language: English
DOI
Text on another site
Keywords: reading comprehensionвопросы с множественным выборомlarge language model (LLM)automatic distractor generationmultiple-choice questionsdataset translationперевод датасетаавтоматическая генерация неправильных вариантов ответапонимание прочитанного текстабольшая языковая модель
Publication based on the results of:
Second-language acquisition modelling within different frameworks of existing theories on the basis of learner corpora platforms for experiments and computer tools (2023)
Similar publications
Субстантивное число в шугнанском языке: корпусное и экспериментальное исследование
Ronko R., Тимофеева В. Д., Индоиранские языки 2026 Т. 2 № 2 С. 117–146
The paper is devoted to the variable number marking in Shughni nouns and reports the results of corpus and experimental studies. In some languages, including Shughni, the marking of plural nouns is not compulsory. Many factors such as animacy, referentiality, morphological and semantic properties, or syntactic position may determine the presence or absence of number marking. The corpus study investigates ...
Added: October 8, 2026
Вариативность форм настоящего времени глагола быть в корпусе говора села Спиридонова Буда (Злынковский район, Брянская область): проблема на стыке грамматики и социолингвистики
Ronko R., Zemicheva S., Труды института русского языка им. В.В. Виноградова 2026 Т. 3 № 49 С. 133–149
In the dialects of the Russian-Belarusian borderland, the verb byt’ ‘to be’ in the present tense is represented by a set of short (e, ё) and full forms (ecть, ёсць, etc.). This article investigates the variability of these forms in the dialect corpus of the village of Spiridonova Buda (Zlynkovskij District, Bryansk Region) and in ...
Added: October 8, 2026
Экспедиции диалектологического кружка НИУ ВШЭ в деревни Западнодвинского района Тверской области
Ronko R., RHEMA. РЕМА 2026 № 1 С. 9–20
In the introductory article to the thematic volume devoted to grammatical research on the speech of several villages, the main linguistic-geographical information about the dialect is presented, along with a description of the key methodological principles for working with the material. The dialects under study belong to the group of the Pskov dialects and border ...
Added: October 8, 2026
Рецепция философии Ф.В.Й. Шеллинга в поздних работах А.Ф. Лосева
Rezvykh P. V., Вестник Российского университета дружбы народов. Серия: Философия 2026 Т. 30 № 1 С. 127–143
The issue delivers an analysis of schellingian motives in the later works by A.F. Losev. The author brings information about Losev’s retrospective evaluation of Schelling’s influence on the formation of his own philosophical approach, it’s specific character and it’s grade. The author shows that all the main investigation projects of the later Losev which have ...
Added: October 7, 2026
Выставка крокодилов в Пассаже в 1864 году: к истории замысла повести «Крокодил» и заметки «Нечто личное» Ф. М. Достоевского
Pershkina A., Slovĕne 2026 Т. 15 № 1 С. 174–192
The article newly presents archival materials from the Directorate of Imperial Theaters relative to the organization of a crocodile exhibition held in St. Petersburg Passage department store in the autumn of 1864, along with newspaper reports on the event. The close analysis of specific articles demonstrates that this exhibition served as the direct inspiration and ...
Added: October 5, 2026
Nomos e poikilía dell'Essere. Genesi metafisica del molteplice. Una deduzione schellinghiana
Dezi A., Sesto San Giovanni: Mimesis, 2026.
This book is a metaphysical inquiry into the nature and origin of the manifold. Taking Schelling’s Erlangen lectures of 1820–1821 as its point of departure and hermeneutic horizon, it develops an interpretation of Nomos that transcends its political, juridical, and social determinations. Nomos thereby emerges as the ontological principle of an infinite, variegated multiplicity of ...
Added: October 2, 2026
Файзов, Д. оии: стихи 2024-2026 годов / Данил Файзов; предисл. М. Павловца. М.: Новое литературное обозрение, 2026. — 160 с. (Серия «Новая поэзия»)
Файзов Д., М.: Новое литературное обозрение, 2026.
Новая книга Д. Файзова пронизана голосами других поэтов. Речь не о расхожей постмодернистской центонности. Это микроинтертекстуальность на уровне конкретных слов и приемов, коллаж осколков чужих голосов, однако таких осколков, которые сохраняют способность отсылать к целому и к тому же вписаны — вклеены на манер аппликации — в узнаваемую интонацию самого поэта. Данил Файзов—поэт, культуртрегер, редактор. ...
Added: October 2, 2026
Manuscript tradition of Kallistos Angelikoudes’ Hesychastic consolation: myths and realities of a fluid text
Vinogradov A., Rodionov O., Travaux et Memoires 2026 Vol. 30 No. 2 P. 1075–1089
This article examines the manuscript tradition of Kallistos Angelikoudes’ Hesychastic consolation, a fluid and complex corpus of thirty Logoi preserved primarily in Vat. gr. 736. The authors survey all known witnesses—including the autograph Barb. gr. 420, the Lond. Arund. 520, and several later copies—and analyze their textual relationships, variant readings, and codicological features. They demonstrate ...
Added: October 1, 2026
Какой национальности лицо? Семантика и прагматика одной конструкции
Громенко Е. С., Krongauz M., Антропологический форум 2026 № 70 С. 163–183
The paper describes the construction litso takoy-to natsionalnosti ‘a person of such-and-such an ethnicity’ in Russian with its semantic and pragmatic characteristics. The purpose of the study is to trace the appearance of the construction and analyse how it is used and how it is perceived. Three expressions representing this construction are considered in more ...
Added: September 30, 2026
Kildin Saami negation is not where you would expect it
Поцелуев В. А., Ural-Altaic Studies 2026 Vol. 62 No. 3 P. 62–81
In recent years, the researchers working on Uralic languages within Minimalism seem to have built a consensus on how to define the position of negation in any language with a negative verb. The main and usually the only argument is based on the work [Mitchell 2006]. The argument is as follows: if the negative verb expresses tense, ...
Added: September 30, 2026
Наука о языке - Новые горизонты. Коллективная монография к юбилеям доктора филологических наук, профессора Ольги Ивановны Максименко, доктора филологических наук, профессора Георгия Теймуразовича Хухуни
Государственный университет просвещения, 2026.
Настоящее издание представляет собой коллективный труд, посвящённый юбилеям основателей-руководителей научных школ лингвистического факультета Государственного университета просвещения Максименко Ольги Ивановны, доктора филологических наук, профессора, профессора кафедры теории языка, англистики и прикладной лингвистики, Председателя диссертационного совета 72.2.020.05 и Хухуни Георгия Теймуразовича, доктора филологических наук, профессора, профессора кафедры теории языка, англистики и прикладной лингвистики. Цель издания – обобщение современных достижений в ...
Added: September 29, 2026
Postnominal numerals are over-specified in referential communication: Evidence from Thai in comparison to Russian
Zevakhina N., Щипкова А. А., Chinkova A., Voprosy Jazykoznanija 2025 No. 2 P. 105–122
This paper presents experimental evidence for the over-specification (redundant use) of Thai postnominal numerals in referential communication. The first experiment reveals high rates of over-specification in oral responses that include Thai postnominal and Russian prenominal numerals tested in a contrastive multi-number visual context. The second experiment demonstrates lower but still relatively high rates of over-specification in written responses that ...
Added: September 29, 2026
Bytedance и Open Source - открытые проекты от разработчика TikTok
Silakov D., Системный администратор 2026 С. 84–89
Social media users rarely think about what lies behind the beautiful facade of activity feeds, teeming with photos and video stories. However, the widespread popularity of such platforms generates a huge amount of all sorts of content that needs to be stored, processed quickly, and displayed, and in the era of AI, it also needs ...
Added: September 28, 2026
FROM LECTURER TO CHATBOT: EVALUATING A HYBRID TEACHING MODEL
Ravedovskaya U., Didenko A., Филатова А. А., Journal of Teaching English for Specific and Academic Purposes 2025 Vol. 13 No. 3 P. 529–538
Emergence of generative artificial intelligence (GenAI) is revolutionizing teaching and evaluation in higher education. Early adopters have already demonstrated how large language model (LLM) conversational agents can serve as on-demand tutors, yet empirical evidence regarding their effectiveness in facilitating conceptual learning in non-computational domains is scant. Based on constructivist learning theory, this paper presents a ...
Added: September 18, 2026
Автоматизированное формирование журналов событий на основе неструктурированных Интернет-источников для задач анализа процессов
Воронова К. Д., Lyadova L. N., Proceedings of the Institute for System Programming of the RAS 2026 Vol. 38 No. 4 P. 153–170
Title: Automated Event Logs Generation Based on Unstructured Internet Sources for Process Analysis Tasks Abstract. This paper presents an approach to automated structuring event-related information extracted from unstructured textual Internet sources for process mining tasks. In many practical cases, information on events associated with various processes is not presented in the form of ready-made event logs, but is ...
Added: September 14, 2026
Benchmarking DNA large language models on quadruplexes
Cherednichenko O., Herbert A., Poptsova M., Computational and Structural Biotechnology Journal 2025 Vol. 27 P. 992–1000
Large language models (LLMs) in genomics have successfully predicted various functional genomic elements. While their performance is typically evaluated using genomic benchmark datasets, it remains unclear which LLM is best suited for specific downstream tasks, particularly for generating whole-genome annotations. Current LLMs in genomics fall into three main categories: transformer-based models, long convolution-based models, and state-space models ...
Added: June 19, 2026
Pre-trained LLMs Meet Sequential Recommenders: Efficient User-Centric Knowledge Distillation
Severin N., Kartushov D., Urzhumov V. et al., , in: Advances in Information Retrieval: 48th European Conference on Information Retrieval, ECIR 2026, Delft, The Netherlands, March 29 – April 2, 2026, Proceedings, Part II. (LNCS, volume 16484).: Cham: Springer Publishing Company, 2026. P. 508–517.
Sequential recommender systems have achieved significant success in modeling temporal user behavior but remain limited in cap-turing rich user semantics beyond interaction patterns. Large Language Models (LLMs) present opportunities to enhance user understanding with their reasoning capabilities, yet existing integration approaches cre-ate prohibitive inference costs in real time. To address these limitations, we present a ...
Added: June 18, 2026
ESQA: Event Sequences Question Answering
Abdullaeva I., Karpukhin I., Filatov A. et al., IEEE Access 2026 Vol. 14 P. 59390–59408
Event sequences, a specialized type of tabular data annotated with timestamps, are prevalent across practical domains such as finance, retail, social networks, and healthcare. Despite the importance of event sequence modeling and analysis, there has been little effort to adapt Large Language Models (LLMs) to this domain. In this paper, we propose a novel solution ...
Added: June 16, 2026
Рефакторинг исходного кода на основе LLM и расширения UML
Karavaeva E., Kuligin L., Rezunik L. et al., Труды Института системного программирования РАН 2026 Т. 38 № 3 С. 67–94
В статье представлен метод рефакторинга исходного кода на основе интеграции большой языковой модели (LLM) и расширенной UML-модели программного кода. Предложенный подход позволяет выявлять проблемные участки кода с использованием функций тревожности и структурных метрик классов, а затем выполнять автоматизированный рефакторинг. Ключевой особенностью метода является использование LLM для генерации формальных спецификаций на языке OCL (Object Constraint Language), ...
Added: May 24, 2026
Large Language Model-Based Automated Item Generation in STEM Assesements: Historical Mapping and a Scoping Review of Empirical Studies
JOURNAL OF EDUCATIONAL TECHNOLOGY DEVELOPMENT AND EXCHANGE 2026 Vol. 19 No. 2 P. 141–169
Educational assessments, from low-stakes classroom tests to high-stakes national examinations, require item pools that are valid, fair, and secure. Automated Item Generation (AIG) aims to efficiently produce large pools of calibrated test items. This paper adopts a two-part design: (1) a brief historical mapping situating LLM-based AIG within the broader AIG trajectory; and (2) a ...
Added: May 5, 2026
LoRA meets Riemannion: Muon Optimizer for Parametrization-independent Low-Rank Adapters
Vladimir Bogachev, Aletov V., Alexander Molozhavenko et al., , in: The Fourteenth International Conference on Learning Representations (ICLR 2026).: ICLR, 2026. Ch. 20503 P. 1–26.
This work presents a novel, fully Riemannian framework for Low-Rank Adaptation (LoRA) that geometrically treats low-rank adapters by optimizing them directly on the fixed-rank manifold. This formulation eliminates the parametrization ambiguity present in standard Euclidean optimizers. Our framework integrates three key components to achieve this: (1) we derive Riemannion, a new Riemannian optimizer on the fixed-rank ...
Added: April 29, 2026
Bridging the Semantic Gap in Metadata Management using Large Language Models
Сулейкин А. С., Сорокина В., Пятецкий В. Е., , in: 2025 7th International Conference on Control Systems, Mathematical Modeling, Automation and Energy Efficiency.: [б.и.], 2025. P. 748–753.
Effective metadata management is fundamental to data governance, ensuring that data assets are discoverable, understandable, and usable across the enterprise. However, traditional metadata systems often remain purely technical, describing structures without conveying business meaning. This disconnect — known as the semantic gap — limits the interpretability and value of metadata for business users. To address ...
Added: April 17, 2026
Разработка и интеграция AI-ассистента в систему управления обучением.
Karavaeva E., Vasilevsky V., Ланин Г. М. et al., Труды Института системного программирования РАН 2025 Т. 37 № 4 С. 175–190
The ongoing digitalization of education requires new ways of presenting information and attention retention mechanisms. The aim of the presented work is to propose a solution for implementing a large language model, which will interactively generate prompts of different types, within an e-learning course on programming. The main approaches are the analysis of existing relatively ...
Added: December 25, 2025
Prediction of protein-protein interactions using point transformer and spherical Convex Hull graphs
David Arteaga, Poptsova M., Computational and Structural Biotechnology Journal 2026 Vol. 31 P. 82–93
Accurate predictions and large-scale identification of protein-protein interactions (PPIs) are crucial for understanding their inherent biological mechanisms and protein functions in virtually all biological processes. Nowadays, graph-based deep learning models have made significant contributions in modeling proteins with physicochemical and geometric features. However, most of these models rely on conventional graph construction methods, such as ...
Added: December 22, 2025
  • About
  • About
  • Key Figures & Facts
  • Sustainability at HSE University
  • Faculties & Departments
  • International Partnerships
  • Faculty & Staff
  • HSE Buildings
  • HSE University for Persons with Disabilities
  • Public Enquiries
  • Studies
  • Admissions
  • Programme Catalogue
  • Undergraduate
  • Graduate
  • Exchange Programmes
  • Summer University
  • Summer Schools
  • Semester in Moscow
  • Business Internship
  • Research
  • International Laboratories
  • Research Centres
  • Research Projects
  • Monitoring Studies
  • Conferences & Seminars
  • Academic Jobs
  • Yasin (April) International Academic Conference on Economic and Social Development
  • Media & Resources
  • Publications by staff
  • HSE Journals
  • Publishing House
  • iq.hse.ru: commentary by HSE experts
  • Library
  • Economic & Social Data Archive
  • Video
  • HSE Repository of Socio-Economic Information
  • HSE1993–2026
  • Contacts
  • Copyright
  • Privacy Policy
  • Site Map
Edit