?
О частотном словаре Национального корпуса русского языка
.
In book
Гродно: ГрГУ, 2007.
Kirina M., Лукьянчикова А. С., В кн.: Язык в эпоху цифровых трансформаций и развития искусственного интеллекта : Сборник научных статей по итогам II Международной научной конференции Минск, 23–24 октября 2025 г.: Мн.: БГУИЯ, 2025. С. 74–85.
В статье рассматриваются характерные особенности гороскопических текстов как части астрологического дискурса. Материалом исследования выступает представительная выборка ежедневных предсказаний на русском языке, опубликованных в открытых группах социальной сети «ВКонтакте», суммарным объемом 1185425 словоупотреблений. С использованием методов корпусной и компьютерной лингвистики анализируются содержательные лексические единицы – как общие, так и отличительные для каждого знака зодиака (в сопоставлении ...
Added: February 28, 2026
Rakhilina E. V., Вестник Российской академии наук 2024 Т. 94 № 9 С. 795–803
Статья посвящена проекту создания Национального корпуса русского языка (НКРЯ) – мощной справочно-информационной системы по русскому языку, которая была разработана консорциумом организаций РАН с участием компании “Яндекс”. Описаны история создания Корпуса, основной его функционал и пути совершенствования, а также наиболее технологичные подкорпуса – поэтический, параллельный, мультимедийный; приведены примеры их работы. Особое внимание уделено последним разработкам, которые ...
Added: February 25, 2025
Kirina M., Лукьянчикова А. С., В кн.: Восьмая Калининградская школа по гуманитарной информатике : сборник докладов. Калининград, 12–14 декабря 2024 года [Электронный ресурс]: научное электронное издание.: Калининград: Смартбукс, 2024. С. 69–73.
В статье рассматриваются лингвостатистические показатели прямой речи литературных персонажей в динамике по историческим периодам. Сопоставляются лексические и морфологические особенности прямой речи и устной речи, представленной в Устном корпусе в составе Национального корпуса русского языка. Материалом исследования стала выборка из 648 рассказов, включенных в Корпус русского рассказа XX века. Объем прямой речи составил 529289 словоупотреблений. На ...
Added: November 29, 2024
Аванесян Н. Л., Зенькова В. В., Chepovskiy A. et al., Успехи кибернетики 2023 Т. 4 № 2 С. 33–39
In this paper the authors describe the methodology for the statistical analysis of texts in the network of Telegram channels based on comparison of automatically generated frequency dictionaries by methods of correlation analysis. Coefficients of pairwise rank correlation are considered for comparing the frequency characteristics of texts in natural language. The method is proposed to ...
Added: July 19, 2023
Баркова Л. А., Труды института русского языка им. В.В. Виноградова 2022 № 3(33) С. 59–95
This article analyzes stanzaic, metric, and rhyming structure in N. E. Gorbanevskaya’s poems, and shows their specifi c features by comparing them with other poems, written in the second half of the 20th century and in the beginning of the 21st century. The source of the present study is the data from the Poetic sub-corpus ...
Added: October 21, 2022
Talalakina E., Stukal D., Kamrotov M., Modern Language Journal 2020 Vol. 104 No. 3 P. 618–646
To date, attempts at empirically validating a construct of academic vocabulary in the form of a frequency list in languages other than English remain conspicuously absent in peer‐reviewed journals. This study aims to close this gap by using Russian as a case study to develop an academic vocabulary list and prove its viability through a ...
Added: August 31, 2020
Orekhov B., Савчук С. О., Труды института русского языка им. В.В. Виноградова 2019 № 21 С. 61–82
В настоящей статье рассмотрено несколько вопросов, связанных с разработкой и использованием акцентологического корпуса в качестве инструмента для исследования ударения: состав и структура корпуса, текущее состояние, перспективы развития, пополнение новым материалом. Особое внимание уделено подкорпусу наивной поэзии в составе акцентологического корпуса как источнику акцентологических данных. Возможности этого ресурса, его эффективное использование проверены на нескольких участках акцентологической ...
Added: March 25, 2020
Lyashevskaya O., Журавлева А. А., В кн.: VII Международные Бодуэновские чтения: Международная конференция И.А. Бодуэн де Куртенэ и мировая лингвистика.: Каз.: Казанский (Приволжский) федеральный университет, 2019.
В статье анализируется смешанная адъективно-генитивная посессивная конструкция в контексте ее представления в синтаксическом формализме Universal Dependencies. Исследование выполнено на материалах частотных синтаксических баз данных поэтического и старорусского корпусов НКРЯ. ...
Added: December 15, 2019
Lyashevskaya O., , in: Computational Linguistics and Intellectual TechnologiesIssue 18.: M.: Russian State University for the Humanitie, 2019. P. 422–434.
The paper discusses the standardization efforts to create a morphological standard for the Middle Russian corpus, which is part of the historical collection of the Russian National Corpus (RNC). To meet the needs of different categories of corpus researchers as well as NLP developers, we consider two styles of the morphological annotation (RNC schema and ...
Added: June 12, 2019
Lyashevskaya O., Vlasova E., Litvintseva K. et al., / NRU HSE. Series WP BRP "Linguistics". 2018. No. 77.
A data analysis tool of the Corpus of Russian Poetry (a part of the Russian National Corpus) is designed for quantitative research in various areas of versology and linguistics aspects of poetic texts. The core part, a statistic database of the corpus, includes annotation at the level of texts, verses, words as well as patterns ...
Added: December 13, 2018
Maslov V., Математические заметки 2017 Т. 101 № 4 С. 531–548
В статье с математической точки зрения рассматриваются аналогии между языком и многочастичными системами в термодинамике. Делается попытка введения математического аппарата и технических средств статистической физики в лингвистические описания. В частности, к лингвистическим объектам применяются понятия числа степеней свободы, бозе-конденсата, фазового перехода и др. На основе статистического анализа словаря и лингвостатистических распределений выдвигается гипотеза о фазовом переходе первого рода от семиотической ...
Added: October 28, 2018
Bonch-Osmolovskaya A. A., Шаги/Steps 2018 № 4 С. 115–146
The paper studies constructions that involve the name of the decade - the twenties, the thirties, the forties etc. – and an adjective in attributive function. The basic assumption is that these constructions reflect the mnemonic pattern of each of the decade from the Soviet and PostSoviet history, the analysis of the constructions therefore is a clue to ...
Added: April 15, 2018
Гаврилова Т. С., Шалганова Т. А., Lyashevskaya O., Вестник Православного Свято-Тихоновского гуманитарного университета. Серия 3: Филология 2017 Т. 51 С. 11–20
The highly unstable orthography of the Middle Russian texts poses a challenge for their automatic processing. The Middle Russian subcorpus of the Russian National Corpus (RNC) includes documents written mainly between 1400 and 1700, when the variation in spelling was still a norm. The task of lexico-grammatical analysis is to assign a dictionary form (lemma), ...
Added: December 14, 2016
Гаврилова Т. С., Шалганова Т. А., Lyashevskaya O., Вестник Православного Свято-Тихоновского гуманитарного университета. Серия 3: Филология 2016 Т. 47 № 2 С. 7–25
The paper discusses two approaches to the automatic lexico-grammatical tagging of the Middle Russian texts (1400–1700), included in the Russian National Corpus (RNC). The task is to assign each token a part of speech label, a tuple of grammatical features, and a lemma (without disambiguation). Middle Russian combines, on the one hand, features of ...
Added: December 14, 2016
Levinzon A. I., Труды института русского языка им. В.В. Виноградова 2015 № 6 С. 641–658
До сих пор на уроках русского языка в российской школе практически не используются электронные корпуса. Цель статьи — продемонстрировать возможности НКРЯ как инструмента эффективной
работы с детьми. Мы анализируем как основные достоинства различных методов корпусной педагогики, так и сложности, которые
предстоит преодолеть учителю, выбравшему, например, метод обучения на основе анализа данных.
Ключевые слова: корпусная педагогика, Национальный корпус
русского языка, ...
Added: March 14, 2016
Bonch-Osmolovskaya A. A., Труды института русского языка им. В.В. Виноградова 2015 Т. 4 № 6 С. 605–641
The goal of the study is to show links between lexical and social diachronic change. The study is conducted in the culturomics framework (Michel et al 2011). In contrast to the Big data approach the study promotes the idea of medium data, i.e. amount of data which allows both to make quantitative and qualitative analysis.The ...
Added: March 14, 2016
Daniel M., , in: Компьютерная лингвистика и интеллектуальные технологии. По материалам ежегодной Международной конференции "Диалог" (2015).: М.: Изд-во РГГУ, 2015. P. 95–103.
The paper discusses the present stage of the evolution of the initial [n]/[j] stem alternation in Russian third person pronouns. After providing a short overview of the origins of the forms, I focus on their category status, discuss Zalizniak’s ‘adpositionality’ in some detail, and then proceed to considering the cases where the ‘n’-forms are induced ...
Added: October 9, 2015
Chepovskiy A., В кн.: Труды Международной научной конференции по физико-технической информатике (CPT2014).: М., Протвино: Институт физико-технической информатики, 2015. С. 120–124.
В статье описана модель изменения словаря естественного языка, отличная от классической модели глоттохронологии. Показана разная скорость убывания ядра языка для различных частей речи. Приведены зависимости для русского языка. ...
Added: July 12, 2015
Lyashevskaya O., М.: Языки славянской культуры, 2016.
Corpus linguistics can be broadly defined in terms of two partially overlapping research dimensions . On the one hand, corpus linguistics is knowledge of how to compile and annotate linguistic corpora. On the other hand, corpus linguistics is a family of qualitative and quantitative methods of language study based on corpus data. The book presents ...
Added: March 26, 2015
Борисенко Н. О., Zhirnova V., Sibirtseva V., В кн.: Язык в различных сферах коммуникации: материалы международной научной конференции.: Чита: Забайкальский государственный университет, 2014. С. 177–181.
The article deals with the derivational models of nouns presented in the textbooks of Russian as a
foreign language “Poekhali -2!”. The article gives an analysis of nouns with suffixes which have meaning
of “face value” (-tel; -chick; -EC; -nick; -anin; -in) in comparison with the data of Russian National
Corpus. The text gives a valuable information on ...
Added: December 1, 2014
Karpov N., Vitugin F., Baranova J., , in: Analysis of Images, Social Networks and TextsVol. 436: 3rd International Conference on Analysis of Images, Social networks, and Texts.: NY: Springer, 2014. Ch. 436 P. 91–100.
In an effort to make reading more accessible, an automated readability formula can help students to retrieve appropriate material for their language level. This study attempts to discover and analyze a set of possible features that can be used for single-sentence readability prediction in Russian. We test the influence of syntactic features on predictability of ...
Added: November 28, 2014
Митрофанова О. А., Lyashevskaya O., Грачкова М. А. et al., В кн.: Структурная и прикладная лингвистикаВып. 9.: СПб.: Издательство СПбГУ, 2012. С. 159–175.
The research project reported in this paper aims at automatic extraction of
linguistic information from contexts in the Russian National Corpus (RNC) and its subsequent use in building a comprehensive lexicographic resource – the Index of Russian lexical constructions. The proposed approach implies automatic context classification intended for word sense disambiguation (WSD) and construction identification (CxI). ...
Added: February 25, 2014
Sibirtseva V., Rocznik Instytutu Polsko-Rossyjskiego 2013 № 2 (5) С. 98–110
Philological research, especially in the field of literature, is usually considered a "thing-in-itself"; the intrinsic value of this phenomenon involves extremely intuitive, creative, "human-readable" analysis. Meanwhile, modern variety of computer programs (semantic text referentors, tag clouds, concordansers, etc.), created also for the humanities, such as sociology, psychology, management, cannot but draw a philologist’s attention. The ...
Added: February 16, 2014