• A
  • A
  • A
  • АБВ
  • АБВ
  • АБВ
  • A
  • A
  • A
  • A
  • A
Обычная версия сайта
  • RU
  • EN
  • HSE University
  • Publications
  • Book chapter
  • RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs
  • RU
  • EN
Расширенный поиск
Высшая школа экономики
Национальный исследовательский университет
Priority areas
  • business informatics
  • economics
  • engineering science
  • humanitarian
  • IT and mathematics
  • law
  • management
  • mathematics
  • sociology
  • state and public administration
by year
  • 2028
  • 2027
  • 2026
  • 2025
  • 2024
  • 2023
  • 2022
  • 2021
  • 2020
  • 2019
  • 2018
  • 2017
  • 2016
  • 2015
  • 2014
  • 2013
  • 2012
  • 2011
  • 2010
  • 2009
  • 2008
  • 2007
  • 2006
  • 2005
  • 2004
  • 2003
  • 2002
  • 2001
  • 2000
  • 1999
  • 1998
  • 1997
  • 1996
  • 1995
  • 1994
  • 1993
  • 1992
  • 1991
  • 1990
  • 1989
  • 1988
  • 1987
  • 1986
  • 1985
  • 1984
  • 1983
  • 1982
  • 1981
  • 1980
  • 1979
  • 1978
  • 1977
  • 1976
  • 1975
  • 1974
  • 1973
  • 1972
  • 1971
  • 1970
  • 1969
  • 1968
  • 1967
  • 1966
  • 1965
  • 1964
  • 1963
  • 1958
  • More
Subject
News
September 11, 2026
How to Assess Students Knowledge in the Age of AI
A researcher at HSE University has proposed a flowchart to help lecturers decide how to assess students who use artificial intelligence. It shows where the use of AI should be restricted and where it can be incorporated into the learning process. The article has been published in IT Professional.
September 9, 2026
‘Balkan Hospitality Opens Doors: Studying Dialects on the Verge of Extinction
You cannot study spoken dialects from books. Instead, you need to go to a village, seek out its elders, and earn the trust of local residents before you can record hours of spontaneous stories. This is how Natalia Muravleva, Associate Professor at the Faculty of Humanities, conducts her research. Her internship in Serbia continued her long-standing study of dialects spoken by Macedonian settlers. In this interview, she discusses how diaspora cultural centres help researchers reach informants, why native speakers need to be interviewed only in their own language (otherwise, as she puts it, they may 'break'), and how a single field season helped her finalise her monograph. She also shares warm memories of autumn in Belgrade and of colleagues with whom grammar can be discussed in three languages at once.
September 9, 2026
Scientists Train Neural Network to Generate Process Plans from 3D Models
Researchers at the HSE FCS AI and Digital Science Institute have developed CAD2TechSpec, a framework that converts 3D models of mechanical parts into machining process plans—step-by-step instructions for machine tools. The solution aims to reduce the time required for the design and preparation of technical process documentation in mechanical engineering, aircraft manufacturing, and other high-tech industries. The study findings have been published in PeerJ Computer Science.

 

Have you spotted a typo?
Highlight it, click Ctrl+Enter and send us a message. Thank you for your help!

Publications
  • Books
  • Articles
  • Chapters of books
  • Working papers
  • Report a publication
  • Research at HSE

?

RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs

P. 9268–9299.
Taktasheva E., Bazhukov M., Koncha K., Fenogenova A., Artemova E., Mikhailov V.

Minimal pairs are a well-established approach to evaluating the grammatical knowledge of language models. However, existing resources for minimal pairs address a limited number of languages and lack diversity of language-specific grammatical phenomena. This paper introduces the Russian Benchmark of Linguistic Minimal Pairs (RuBLiMP), which includes 45k pairs of sentences that differ in grammaticality and isolate a morphological, syntactic, or semantic phenomenon. In contrast to existing benchmarks of linguistic minimal pairs, RuBLiMP is created by applying linguistic perturbations to automatically annotated sentences from open text corpora and decontaminating test data. We describe the data collection protocol and present the results of evaluating 25 language models in various scenarios. We find that the widely used LMs for Russian are sensitive to morphological and agreement-oriented contrasts, but fall behind humans on phenomena requiring the understanding of structural relations, negation, transitivity, and tense. RuBLiMP, the codebase, and other materials are publicly available.

Language: English
Full text
DOI
Text on another site
Keywords: Russianminimal pairsbenchmarkbenchmark dataset

In book

Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing
Association for Computational Linguistics, 2024.
Similar publications
Constructions Built of Constructions: Intensified Comparatives
Rakhilina E. V., Zhukova V., Akhapkina Y., Zeitschrift fur Slawistik 2026 Vol. 71 No. 2 P. 259–284
This study examines the most numerous type of construction presented in the Russian Constructicon, namely intensifiers. We studied intensifiers in the context of comparative constructions of inequality like Х bol’še Y‘X is bigger than Y’ –> Х gorazdo bol’še Y‘X is much bigger than Y’. Comparative constructions and intensifiers are usually studied separately since they ...
Added: September 4, 2026
ML²B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation
Trofimova E., Shamina Z., Selifanova M. et al., , in: Proceedings of the Generative Code Intelligence Workshop (GeCoIn 2026), co-located with the 35th International Joint Conference on Artificial Intelligence (IJCAI-ECAI 2026)Vol. 4238.: CEUR-WS.org, 2026.
We introduce ML2B, the first benchmark for evaluating cross-lingual task comprehension in end-to-end ML pipeline generation by large language models. Despite growing global AI adoption, no systematic evaluation exists for ML pipeline generation beyond English task descriptions. ML2B addresses this gap with 35 Kaggle competitions spanning tabular, text, and image domains, translated into 14 languages ...
Added: August 20, 2026
RusLan-M: longitudinal multimedia corpus of early child speech in Russian
Diachkova M., Lelik V., Dorofeeva S. et al., Language Resources and Evaluation 2026 Vol. 60 Article 68
This article presents the Russian Language-Monolingual corpus (RusLan-M, v.1.0), a longitudinal multimedia collection of early child speech from two Russian-speaking monolingual children: Tosya (ages 0;10–3;10, 246 recordings) and Yasha (ages 1;04–3;00, 42 recordings). The corpus consists of approximately 41 h (2,454 min.) of video recordings and 35,386 child utterances, available with transcriptions in the CHAT ...
Added: August 17, 2026
Kinh nghiệm cửa LB Nga trong việc viết sách giáo khoa về dịch thuật tiếng Việt
Britov I., Tyumeneva E., Glazunova S., , in: Ngoại giao, biên phiên dịch và hợp tác quốc tế của Việt Nam trong kỷ nguyên mới.: Hà Nội: Nhà xuất bản Đại học quốc gia Hà Nội, 2026. P. 417–427.
The article analyzes the Russian experience in writing textbooks on translation from Vietnamese into Russian and from Russian into Vietnamese. On the basis of five textbooks written in Russia over the past thirty years, forms and methods of teaching both oral and written translation are revealed. A comparative description of each of the presented textbooks ...
Added: August 11, 2026
Discovering the Potential of Automated Phraseological Interference Error Detection: A Transformer-Based Approach
Kharlamova D. S., Journal of the European Second Language Association 2026 Vol. 10 No. 1 P. 1–16
Formulaic language may help language learners in second language (L2) acquisition. However, interference with the first language (L1) can also cause errors in L2 production. The present paper explores the possibilities of detecting L1 Russian interference errors connected with phraseologisms in English learner texts with a fine-tuned Transformer-based neural network. Across a dataset of 3,600 ...
Added: August 6, 2026
Local Fault-Tolerant Routing in 3D Mesh NoCs using Single-Hop Rollback
Edward R. Rzaev, Aleksandr Y. Romanov, Andrey M. Sukhov, IEEE Access 2026 Vol. 14 P. 112184–112201
This work presents a hierarchy of strictly local fault-tolerant routing algorithms for 3D mesh networks-on-chip, culminating in an algorithm that combines a live-neighbor selection rule with a bounded single-hop rollback mechanism. The proposed algorithms operate exclusively on immediate neighbor information, maintain O(1) per hop complexity, and require no global topology knowledge, additional virtual channels, or ...
Added: July 23, 2026
Систематизация равноправных произносительных вариантов в современном русском языке (на материале орфоэпических словарей)
Zubov V., Вопросы лексикографии 2026 № 40 С. 64–86
The article addresses the problem of selecting and systematizing data for the study of pronunciation variation in contemporary Russian and proposes a solution in the form of a specialized database of codified equivalent pronunciation variants (e.g., simmétriya / simmetríya “symmetry”). The article presents a methodology for identifying, selecting, and organizing such variants into a database. ...
Added: July 23, 2026
RussianSuperGLUE: A Russian Language Understanding Evaluation Benchmark
Shavrina T., Fenogenova A., Emelyanov A. et al., , in: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP).: Association for Computational Linguistics, 2020. P. 4717–4726.
In this paper, we introduce an advanced Russian general language understanding evaluation benchmark – RussianSuperGLUE. Recent advances in the field of universal language models and transformers require the development of a methodology for their broad diagnostics and testing for general intellectual skills - detection of natural language inference, commonsense reasoning, ability to perform simple logical ...
Added: June 14, 2026
Listen, Repeat, Decide: Investigating Pronunciation Variation in Spoken Word Recognition among Russian Speakers
Zubov V., Elena Riekhakaynen, , in: Proceedings of the Workshop on Cognitive Aspects of the Lexicon @ LREC-COLING 2024.: European Language Resources Association (ELRA), 2024. P. 129–132.
Variability is one of the important features of natural speech and a challenge for spoken word recognition models and automatic speech recognition systems. We conducted two preliminary experiments aimed at finding out whether native Russian speakers regard differently certain types of pronunciation variation when the variants are equally possible according to orthoepic norms. In the ...
Added: April 19, 2026
HoTPP benchmark: Are we good at the long horizon events forecasting?
Karpukhin I., Shipilov F., Savchenko A., Neurocomputing 2026 Vol. 672 Article 132771
Forecasting multiple future events within a given time horizon is essential for applications in finance, retail, social networks, and healthcare. This problem is typically addressed using Marked Temporal Point Processes (MTPP), which provide a principled framework for modeling both event timing and event labels. While most existing research focuses on predicting only the next event, forecasting distant future ...
Added: February 25, 2026
Difference in Language Profiles of Children With Autism Spectrum Disorder and Down Syndrome Is Not Driven by Non-Verbal Cognition
Novoselova K., Lopukhina A., Gomozova M. et al., International Journal of Language and Communication Disorders 2026 Vol. 61 No. 1 Article e70177
Background Autism Spectrum Disorder (ASD) and Down syndrome (DS) are among the most common types of neurodevelopmental conditions that have co-occurring language impairments. Usually, non-verbal IQ has been reported as one of the main predictors of language functioning in children with these conditions. Although language abilities of children with ASD and DS have been described in ...
Added: February 6, 2026
Experimental evidence suggests that null complement anaphora in Russian is not reducible to clausal ellipsis
Knyazev M., Folia Linguistica 2026 Vol. 60 No. 1 P. 453–496
Null complement anaphora, NCA (e.g., I suggested the price was too high, and she agreed ∅.), is a long known but poorly understood phenomenon subject to idiosyncratic lexical restrictions. In languages like Russian, however, it is (or appears) productive, with verbs not allowing NCA hard to nd, raising the question whether omission of the clausal argument ...
Added: January 19, 2026
Null and overt subjects in Russian polarity focus: Interactions with ellipsis
Kasenov D., Rudnev P., , in: Экспериментальные исследования языка: материалы конференции 2025.: М.: Наш мир, 2025. P. 50–53.
Added: January 19, 2026
Русский язык и русская культура во Вьетнаме: проблемы обучения и исследования
Britov I., Ханой: Ханойский государственный университет, 2025.
Без аннотации ...
Added: January 18, 2026
ComputAgeBench: Epigenetic Aging Clocks Benchmark
Kriukov D., Efimov E., Kuzmina E. et al., , in: KDD '25: Proceedings of the 31th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. Volume 2.: Association for Computing Machinery (ACM), 2025. P. 5560–5570.
The success of clinical trials of longevity drugs relies heavily on identifying integrative health and aging biomarkers, such as biological age. Epigenetic aging clocks predict the biological age of individuals using their DNA methylation profiles, commonly retrieved from blood samples. However, there is no standardized methodology to validate and compare epigenetic clock models. We propose ComputAgeBench, ...
Added: January 12, 2026
The Russian Rey Auditory Verbal Learning Test (RAVLT): version comparison and normative data for children aged 5–18 years
Buivolova O., Malyutina S., Morozova A. et al., Child Neuropsychology 2026 Vol. 32 No. 3 P. 316–331
The Rey Auditory Verbal Learning Test (RAVLT) is a widely used neuropsychological tool developed for assessing various aspects of verbal memory. We present a RAVLT version for Russian-speaking children, developed in digital form with two sets of materials. The current study aimed to investigate whether the two versions of the Russian RAVLT are equivalent in ...
Added: September 24, 2025
LexiaD, the first dyslexia-specific Cyrillic font compared to the popular Times New Roman and Roboto fonts when read by adults
Alexeeva S. V., Zubov V., Nikonova Y., , in: Psychological Applications Conference and Trends (InPACT 2022).: inScience Press, 2022. P. 464–468.
The LexiaD font was developed for Russian-speaking people with reading disorders (dyslexia) (Alexeeva et al., 2020). LexiaD demonstrated an advantage in letter feature extraction and information integration over other modern Cyrillic fonts (PT Sans and PT Serif) while reading by primary school dyslexic and non-dyslexic children. However, for dyslexic and non-dyslexic adolescents, the familiar Arial font was more effective ...
Added: September 23, 2025
  • About
  • About
  • Key Figures & Facts
  • Sustainability at HSE University
  • Faculties & Departments
  • International Partnerships
  • Faculty & Staff
  • HSE Buildings
  • HSE University for Persons with Disabilities
  • Public Enquiries
  • Studies
  • Admissions
  • Programme Catalogue
  • Undergraduate
  • Graduate
  • Exchange Programmes
  • Summer University
  • Summer Schools
  • Semester in Moscow
  • Business Internship
  • Research
  • International Laboratories
  • Research Centres
  • Research Projects
  • Monitoring Studies
  • Conferences & Seminars
  • Academic Jobs
  • Yasin (April) International Academic Conference on Economic and Social Development
  • Media & Resources
  • Publications by staff
  • HSE Journals
  • Publishing House
  • iq.hse.ru: commentary by HSE experts
  • Library
  • Economic & Social Data Archive
  • Video
  • HSE Repository of Socio-Economic Information
  • HSE1993–2026
  • Contacts
  • Copyright
  • Privacy Policy
  • Site Map
Edit