?
A Reusable Tagset for the Morphologically Rich Language in Change: a Case of Middle Russian
P. 422–434.
The paper discusses the standardization efforts to create a morphological standard for the Middle Russian corpus, which is part of the historical collection of the Russian National Corpus (RNC). To meet the needs of different categories of corpus researchers as well as NLP developers, we consider two styles of the morphological annotation (RNC schema and Universal Dependencies schema). A number of specifications of the feature list proposed to facilitate data reusability, linking and conversion.
Keywords: Национальный корпус русского языкадревнерусский языклемматизацияRussian National Corpusлексико-грамматическая разметкаuniversal dependenciesMiddle Russianстарорусская письменностьисторические корпусаlemmatizationOld RussianPOS taggingчастеречная разметкаfull morphology taggingtagsethistorical corporaтагсет
Publication based on the results of:
In book
Issue 18. , M.: Russian State University for the Humanitie, 2019.
Ermolova M., Russian Linguistics 2026 Т. 50 Статья 14
The article analyzes the functioning of gerunds in the Russian language of the 17th century. Basedon the analysis of contexts that are absent in modernRussian, itisconcludedthatinthe 17th century the gerund lost the absolute temporal meaning it once had, acquiring a relative meaning depending on the tense of the main predicate, while remaining, at the same ...
Added: July 4, 2026
Глазкова А. В., Смаль И. В., Lyashevskaya O. et al., Доклады Российской академии наук. Математика, информатика, процессы управления (ранее - Доклады Академии Наук. Математика) 2025 Т. 527 С. 146–155
This paper presents a study on the effectiveness of discriminative methods for abbreviation lemmatization in Russian texts. Unlike generative approaches, discriminative models select the optimal lemma from a fixed set of candidates, eliminating the risk of generating grammatically incorrect word forms. For the first time in Russian language processing, we conduct a comprehensive analysis of ...
Added: March 10, 2026
Afanasev I., Glazkova A., Lyashevskaya O. et al., , in: Proceedings of the 10th Workshop on Slavic Natural Language Processing (Slavic NLP 2025).: Association for Computational Linguistics, 2025. P. 157–170.
Pre-trained language models have significantly advanced natural language processing (NLP), particularly in analyzing languages with complex morphological structures. This study addresses lemmatization for the Russian language, the errors in which can critically affect the performance of information retrieval, question answering, and other tasks. We present the results of experiments on generative lemmatization using pre-trained language ...
Added: March 10, 2026
Glazkova A., Lyashevskaya O., Morozov D. et al., Journal of Mathematical Sciences 2025 Vol. 546 P. 32–47
This paper addresses the task of lemmatizing abbreviations in the Russian language. Abbreviation lemmatization is particularly challenging, as it involves not only transforming a word into its normal form but also correctly expanding the abbreviation. We explore two approaches to this task, both leveraging large pretrained language models. The first approach is generative, where the ...
Added: March 10, 2026
Ronko R., Wiemer B., , in: Encyclopedia of Slavic Languages and Linguistics Online.: Brill, 2020.
The nominative object describes a clause type in which the object of a transitive verb takes nominative morphology, and this coding is not conditioned by voice operations. It is a salient property in regions in which Slavic varieties have been in contact with Finnic- and/or Baltic-speaking population, i.e., in the eastern part of the Circum-Baltic ...
Added: December 19, 2025
Shumen: INCOMA Ltd, 2025.
This paper introduces a rule-based lemmatization and word embedding pipeline for the endangered Bartangi language, part of the Pamiri language group. The system combines a manually constructed lemma dictionary with morphological suffix rules to improve linguistic consistency in low-resource settings. The results demonstrate enhanced lemmatization accuracy and higher-quality embeddings for downstream NLP tasks. The work ...
Added: October 20, 2025
Anna A. Fitiskina, Russian linguistics 2025 Vol. 49 Article 4
This paper aims to demonstrate that the Old East Slavic pronoun iže, traditionally considered a loanword from Old Church Slavonic and a marker of literacy, was in fact also widely used in secular texts of the earliest period and that its usage there differed considerably from that found in Old East Slavic church-oriented literature. The ...
Added: September 26, 2025
Mylnikova A., Mylnikov L., Научно-техническая информация. Серия 2: Информационные процессы и системы 2025 № 7 С. 32–44
Рассмотрена модель использования скелетных структур на базе синтаксической разметки для предобработки корпусов текстов перед передачей в нейросетевые модели машинного перевода с целью повышения качества их работы, реализованная с помощью частеречной и синтаксической разметок корпусов текстов, использующих языковую модель, с использованием сети BERT и набора правил. Описана подготовка данных для обучения и предложены способы повышения эффективности ...
Added: September 22, 2025
Gippius A., Вопросы языкознания 2025 № 4 С. 7–41
This article contains a preliminary publication of 30 birchbark letters found during the 2024 archaeological season at the Troitsky excavation in Veliky Novgorod. The vast majority of the published texts date back to the 12th century. Most important in historical and philological terms are the following items: a letter mentioning a military campaign and related ...
Added: September 21, 2025
Rakhilina E. V., Вестник Российской академии наук 2024 Т. 94 № 9 С. 795–803
Статья посвящена проекту создания Национального корпуса русского языка (НКРЯ) – мощной справочно-информационной системы по русскому языку, которая была разработана консорциумом организаций РАН с участием компании “Яндекс”. Описаны история создания Корпуса, основной его функционал и пути совершенствования, а также наиболее технологичные подкорпуса – поэтический, параллельный, мультимедийный; приведены примеры их работы. Особое внимание уделено последним разработкам, которые ...
Added: February 25, 2025
Plungian V., Вестник Российской академии наук 2024 Т. 94 № 9 С. 787–794
Даётся общее представление о корпусной лингвистике, её истории, методах и влиянии на современные представления об изучении языка, которое обычно обозначается как “корпусная революция”. ...
Added: December 16, 2024
Gippius A., Вопросы языкознания 2024 № 4 С. 7–26
The article contains a preliminary publication of nineteen birchbark letters found during the archaeological season of 2023 in Veliky Novgorod (Nos. 1158–1172) and Staraya Russa (Nos. 55–58). The published documents date back to the 12th— early 16th centuries. From the historical point of view, three 14th-century documents are of the greatest value: No. 1164 is ...
Added: September 7, 2024
Фитискина А. А., В кн.: От сорочка к Олекше: Сборник статьей к 60-летию А. А. Гиппиуса.: М.: РАНХиГС, 2023.
This article is devoted to the history of the word promuzgy (nom. pl.), which is known from Kirik the Novgorodian’s Teaching, a 12th-century treatise on mathematics and the calendar. The word is often considered a hapax, although it is in fact also found in the Cyrillic text of the Boyana Palimpsest and in the Pandects ...
Added: May 15, 2024
Gippius A., Вопросы языкознания 2023 № 5 С. 7–28
: The article contains a preliminary publication of twelve birchbark letters of the twelfth— first half of the fifteenth century, found in the archaeological season of 2022 in Veliky Novgorod (Nos. 1146– 1157), and letters Nos. 52 and 53 from Staraya Russa. Letters Nos. 1142 and 1143 from the excavations of 2021, which were not included ...
Added: February 13, 2024