?
Применение меры tf-idf и меры странности для выделения ключевых слов при классификации текстов научных статей
С. 42–42.
Козлова Е. С., Romanov A.
In book
Сумы: СумДу, 2016.
Chernyavskiy A., Ilvovsky D., Nakov P., , in: CLEF 2021 Working Notes.: CEUR Workshop Proceedings, 2021. P. 484–493.
We describe our system for the CLEF 2021 CheckThat! Lab Task 2 Subtask A on detecting previously fact-checked claims. We developed a pipeline using TF.IDF, sentence-BERT fine-tuned on the training data, and reranking using LambdaMART and the predicted similarity scores and positions in the ranked list as features. We examined the quality of each model ...
Added: May 9, 2024
Remnev N., , in: 2019 International Conference on Data Mining Workshops (ICDMW).: IEEE, 2019. P. 1–7.
The task of recognizing the author’s native language based on a text (Native Language Identification - NLI) is the task of automatically recognizing native language (L1) based on texts written in a language that is not native to the author. The NLI task was studied in detail for the English language, and two shared tasks ...
Added: October 18, 2021
Remnev N., , in: Компьютерная лингвистика и интеллектуальные технологии: по материалам ежегодной международной конференции «Диалог» (Москва, 17–20 июня 2020 г.)Issue 19(26): дополнительный том.: -, 2020. P. 1123–1133.
The task of recognizing the author’s native (Native Language Identification—NLI) language based on a texts, written in a language that is non-native to the author—is the task of automatically recognizing native language (L1). The NLI task was studied in detail for the English language, and two shared tasks were conducted in 2013 and 2017, where ...
Added: October 18, 2021
Romanov A., Lomotin K.E., Kozlova E.S., , in: Supplementary Proceedings of the Sixth International Conference on Analysis of Images, Social Networks and Texts (AIST-SUP 2017), Moscow, Russia, July 27-29, 2017Vol. 1975.: Aachen: CEUR-WS.org, 2017. P. 122–133.
This research examines the problems of automatic scientific articles classification according to Universal Decimal Classifier. To reveal the structure of the train data its visualization was obtained using the recursive feature elimination algorithm. Further; the study provides a comparison of TF-IDF and Weirdness – two statistic-based metrics of keyword significance. The most efficient classification methods ...
Added: November 28, 2017
Lomotin K. E., Kozlova E. S., Romanov A., , in: Information Innovative Technologies: Materials of the International scientific–рractical conference.: M.: Association of graduates and employees of AFEA named after prof. Zhukovsky, 2017. P. 359–363.
The research is devoted to studying of applicability of most relevant modern classification methods to the issue of automatic universal decimal classificator code generation for arbitrary scientific article. The next methods are considered as classifiers: artificial neural network, logistic regression, naive Bayesian classifier and metrical ...
Added: July 30, 2017
Romanov A., Ломотин К. Е., Козлова Е. С., Информационные технологии 2017 Т. 23 № 6 С. 418–423
The paper deals with the applicability of modern machine learning methods to the problem of automatic generation of UDC for scientific articles. As the classifiers, such models as artificial neural networks, logistic regression and boosting are considered. Graph algorithms and a prototype software module to generate UDC are designed. ...
Added: July 30, 2017
Ломотин К. Е., Козлова Е. С., Колесниченко А. Л. et al., В кн.: Инновационные, информационные и коммуникационные технологии: сборник трудов XIII Международной научно-практической конференции.: М.: Ассоциация выпускников и сотрудников ВВИА им. проф. Жуковского, 2016. С. 92–95.
In the article an efficiency of the modern classification instruments usage for the task of scientific articles texts rubrication according to the UDC classifier is analyzed. The following means are explored: artificial neural networks, cosine similarity, naive Bayesian classifier, decision trees and random forest. ...
Added: October 29, 2016
Ломотин К. Е., Romanov A., В кн.: Информатика, математика, автоматика: 2016. Материалы научно-технической конференции.: Сумы: СумДу, 2016. С. 43–43.
Использование искусственных нейронных сетей (ИНС) для решения задач классификации позволяет разделить такие сложные классы образов, какими являются темы классификатора УДК. Для проведения исследования нами выбран классификатор гиперплоскостной группы, реализованный в виде многослойного персептрона Розенблатта. ...
Added: June 11, 2016
Romanov A., Lomotin K.E., Kozlova E.S. et al., , in: 2016 International Siberian Conference on Control and Communications (SIBCON). Proceedings.: M.: HSE, 2016. Ch. 543fu4t.
In this work realization of automatic scientific articles classification according to Universal Decimal Classifier is presented. Efficiency of neural networks technologies application for current task is researched, and optimal neural network structure and parameters are offered ...
Added: June 11, 2016
Starykh V., Белоозеров В. Н., Scientific and Technical Information Processing 2010 № 9 С. 25–34
This article describes how to use and outcomes of the thematic Categorization of information and educational resources. It is based on the Universal Decimal Classification, which has international status and mandatory to describe the subject matter of scientific and technical information. During the first stage finishes compiling the categories for thematic subjects of secondary education ...
Added: October 14, 2013