?
Ассоциативная (однородная семантическая) сеть - семантический компонент системы распознавания слитной речи
.
Глазкова А. В., Смаль И. В., Lyashevskaya O. et al., Доклады Российской академии наук. Математика, информатика, процессы управления (ранее - Доклады Академии Наук. Математика) 2025 Т. 527 С. 146–155
This paper presents a study on the effectiveness of discriminative methods for abbreviation lemmatization in Russian texts. Unlike generative approaches, discriminative models select the optimal lemma from a fixed set of candidates, eliminating the risk of generating grammatically incorrect word forms. For the first time in Russian language processing, we conduct a comprehensive analysis of ...
Added: March 10, 2026
Mylnikova A., Mylnikov L., Научно-техническая информация. Серия 2: Информационные процессы и системы 2025 № 7 С. 32–44
Рассмотрена модель использования скелетных структур на базе синтаксической разметки для предобработки корпусов текстов перед передачей в нейросетевые модели машинного перевода с целью повышения качества их работы, реализованная с помощью частеречной и синтаксической разметок корпусов текстов, использующих языковую модель, с использованием сети BERT и набора правил. Описана подготовка данных для обучения и предложены способы повышения эффективности ...
Added: September 22, 2025
Белова П.Е., Юрислингвистика 2023 № 27(38) С. 94–98
Within the framework of linguistic expertise on cases of copyright and related rights infringement experts are increasingly faced with the challenge of comparing several texts and searching for full-text, partial and other (lexical, grammatical, semantic, etc.) coincidences in them, as well as determining the values of these coincidences. Comparing documents manually takes a lot of ...
Added: April 6, 2023
Кусакин И. К., Цурупа А. М., Алмакаев А. В. et al., В кн.: НТИ-2022. Научная информация в современном мире: глобальные вызовы и национальные приоритеты : материалы 10-ой научной конференции с международным участием, посвященной 70-летию ВИНИТИ РАН, Москва, 25–26 октября 2022 года.: М.: ВИНИТИ РАН, 2022. С. 103–109.
This work is devoted to the study of approaches for training BERT-based classifiers of scientific articles to implement the application with the adoption of the best models for use in the infrastructure of the VINITI RAS. For this purpose, the BERT linguistic model was trained on a specialized corpus of scientific texts for subsequent use ...
Added: January 31, 2023
Кусакин И. К., Федорец О. В., Romanov A., Научно-техническая информация. Серия 2: Информационные процессы и системы 2022 Т. 12 С. 6–9
This paper discusses modern approaches to natural language processing and appliance of artificial intelligence technologies in the task of classifying scientific texts in Russian. The report contains an analysis of implementations of text vectorization methods, a description of experiments with training various classifier models: from classical machine learning algorithms to neural network transformer architectures. ...
Added: January 31, 2023
Sergey Smetanin, Mathematics 2022 Vol. 10 No. 16 Article 2947
Policymakers and researchers worldwide are interested in measuring the subjective well-being (SWB) of populations. In recent years, new approaches to measuring SWB have begun to appear, using digital traces as the main source of information, and show potential to overcome the shortcomings of traditional survey-based methods. In this paper, we propose the formal model for ...
Added: August 15, 2022
Smetanin S., PeerJ Computer Science 2022 No. 8 Article e1039
The Russian language is still not as well resourced as English, especially in the field of sentiment analysis of Twitter content. Though several sentiment analysis datasets of tweets in Russia exist, they all are either automatically annotated or manually annotated by one annotator. Thus, there is no inter-annotator agreement, or annotation may be focused on ...
Added: June 29, 2022
Kharlamov A. A., Kulikov A., , in: Neuroinformatics and Semantic Representations: Theory and Applications.: Cambridge Scholars Publishing, 2020. P. 219–231.
В работе показано использование механизма сравнения семантических сетей текстов в задаче диагностики заболеваний с использованием сигнальных сетей. Выявление степени пересечения семантических сетей текстов позволяет говорить о степени их смыслового подобия. Однородная семантическая сеть как множество узлов, связанных дугами, имеет численные характеристики – частоты появления слов, а также пар слов в тексте, которые перенормируются с использованием ...
Added: December 7, 2021
Kharlamov A. A., , in: Neuroinformatics and Semantic Representations: Theory and Applications.: Cambridge Scholars Publishing, 2020. P. 156–167.
На основе представлений об обработке информации в мозге человека [1] реализована технология автоматической смысловой обработки текстов TextAnalyst, позволяющая выявить ключевые понятия текста в их взаимосвязях, реализовать реферирование текстов и их смысловое сравнение (классификацию). Реализованы продукты, использующие функциональность этой технологии: персональный – TextAnalyst, и библиотека COM модулей – TextAnalyst SDK. ...
Added: December 7, 2021
Kazartsev (Evgenii Kazartcev) E., Долженко Д. Ю., Емельянов Н. И., В кн.: International Scientific and Theoretical Conference "Problems of poetry and prosody IX" is dedicated to the 30th anniversary of the Independence of the Republic of Kazakhstan and the 90th Anniversary of the Outstanding Kazakh Poet Mukagali Makatayev 20-21 May 2021 Almaty.: КазНПУ им. Абая, 2021. С. 35–40.
Статья посвящена анализу влияния стихотворного ритма на прозу Б. Л. Пастернака и В. В. Набокова. В результате исследования стихоподобных фрагментов прозы у этих авторов были отмечены существенные отличия в проявлении поэтических тенденций. В процессе эволюции ритмика прозы Пастернака обнаруживает сближение с ритмикой классических образцов русского стиха, в то время как Набоков напротив со временем отходит ...
Added: October 31, 2021
Smetanin S., , in: Компьютерная лингвистика и интеллектуальные технологии: по материалам ежегодной международной конференции «Диалог» (Москва, 17–20 июня 2020 г.)Issue 19(26): дополнительный том.: -, 2020. P. 1149–1159.
Added: November 30, 2020
Byzov A., Социология: методология, методы, математическое моделирование 2019 № 49 С. 131–160
Throughout most of their history, sociologists have sought to study unstructured organic texts: newspaper materials, diaries, memoirs, letters, documents, and, more recently, messages, publications and other texts on various online platforms. This article discusses how modern techniques of text mining can improve classical sociological approaches to the analysis of this type of data. The article ...
Added: December 9, 2019
Malafeev A., Nikolaev K., , in: Analysis of Images, Social Networks and Texts. 8th International Conference, AIST 2019, Kazan, Russia, July 17–19, 2019, Revised Selected Papers. Communications in Computer and Information ScienceVol. 1086.: Springer, 2020. P. 154–159.
In this paper, a deep learning method study is conducted to solve a new multiclass text classification problem, identifying user interests by text messages. We used an original dataset of almost 90 thousand forum text messages, labeled for ten interests. We experimented with different modern neural network architectures: recurrent and convolutional, as well as simpler ...
Added: November 7, 2019
Kharlamov A. A., Информационные технологии 2017 № 1 С. 66–75
Ассоциативная память человека является средой для формирования единого пространства знаний. Рассмотрена обработка текстовой информации как пример процесса обработки человеком информации любой модальности с формированием однородной семантической сети ключевых понятий текста, ранжированных по степени их важности, а также иерархической тематической структуры, характеризующей сложность текста. Результаты текстового анализа наглядно интерпретируются на анализе конкретных текстов с применением аппарата ...
Added: February 20, 2019
Наконечная Е. Т., , in: Proceedings of the 22nd Conference of Open Innovations Association FRUCT.: Jyvaskyla: [б.и.], 2018. P. 361–365.
Статья является продолжением ряда исследований, посвященных изучению ритмики художественной прозы А. С. Пушкина. В работе рассматриваются такие произведения, как «Дубровский», «Пиковая дама», «Капитанская дочка», «Кирджали», «Египетские ночи». Применяется метод отбора «случайных» четырехстопных ямбов. Ритмика стихоподобных фрагментов сравнивается с вероятностно-статистическими моделями распределения стихотворных строк в прозе. В результате анализа случайных стихоподобных фрагментов рассмотрена эволюция ритмики прозы ...
Added: June 12, 2018
Builova N., Научно-техническая информация. Серия 2: Информационные процессы и системы 2018 № 8 С. 34–38
The problem of documents classification by genre was examined in this review. The main characteristics of the text used to recognize the genre of text were highlighted, and the most widely used algorithms of machine learning were described. The methods considered serve for the classification of scientific, technical, journalistic and artistic texts. ...
Added: March 28, 2018