• A
  • A
  • A
  • АБВ
  • АБВ
  • АБВ
  • A
  • A
  • A
  • A
  • A
Обычная версия сайта
  • RU
  • EN
  • HSE University
  • Publications
  • Book chapter
  • Automated Word Sense Frequency Estimation for Russian Nouns
  • RU
  • EN
Расширенный поиск
Высшая школа экономики
Национальный исследовательский университет
Priority areas
  • business informatics
  • economics
  • engineering science
  • humanitarian
  • IT and mathematics
  • law
  • management
  • mathematics
  • sociology
  • state and public administration
by year
  • 2027
  • 2026
  • 2025
  • 2024
  • 2023
  • 2022
  • 2021
  • 2020
  • 2019
  • 2018
  • 2017
  • 2016
  • 2015
  • 2014
  • 2013
  • 2012
  • 2011
  • 2010
  • 2009
  • 2008
  • 2007
  • 2006
  • 2005
  • 2004
  • 2003
  • 2002
  • 2001
  • 2000
  • 1999
  • 1998
  • 1997
  • 1996
  • 1995
  • 1994
  • 1993
  • 1992
  • 1991
  • 1990
  • 1989
  • 1988
  • 1987
  • 1986
  • 1985
  • 1984
  • 1983
  • 1982
  • 1981
  • 1980
  • 1979
  • 1978
  • 1977
  • 1976
  • 1975
  • 1974
  • 1973
  • 1972
  • 1971
  • 1970
  • 1969
  • 1968
  • 1967
  • 1966
  • 1965
  • 1964
  • 1963
  • 1958
  • More
Subject
News
August 25, 2026
Scientists Develop Algorithm for More Reliable Processors in Data Centres
Researchers from HSE MIEM and Samara University have developed the LRF-3D algorithm to automatically bypass idle nodes in three-dimensional networks-on-chip. Thanks to its hierarchical architecture, the algorithm outperforms existing solutions in both speed and path accuracy, improving processor reliability for use in data centres, supercomputers, and AI computing. The source code and test results are publicly available.
August 24, 2026
Researchers Develop Method for Direct Generation of Regulatory DNA
Researchers at HSE University have developed a model for generating promoters and enhancers—DNA sequences that regulate gene activity. The model works directly with DNA nucleotides, without first transforming them into a continuous numerical representation. This solution could be useful for applications in synthetic biology and gene therapy. The study results were presented at the ICLR 2026 Workshop ‘Generative AI in Genomics (Gen^2): Barriers and Frontiers.’
August 21, 2026
Social Integration: At the Crossroads of Knowledge and Values
The International Laboratory for Social Integration Research (ILSIR) at HSE University studies the challenges faced by vulnerable groups and explores ways to help them participate fully in everyday life. To develop effective solutions, the laboratory’s researchers combine cutting-edge methods with practical fieldwork. In this interview with the HSE News Service, Laboratory Head Elena Iarskaia-Smirnova discusses the laboratory’s work.

 

Have you spotted a typo?
Highlight it, click Ctrl+Enter and send us a message. Thank you for your help!

Publications
  • Books
  • Articles
  • Chapters of books
  • Working papers
  • Report a publication
  • Research at HSE

?

Automated Word Sense Frequency Estimation for Russian Nouns

P. 79–94.
Lopukhina A., Лопухин К. А., Носырев Г. В.

According to G. K. Zipf’s observation, there is a strong correlation between word frequency and polysemy. Yet word sense frequency distribution is a neglected area in computational linguistics. Furthermore, the study of sense frequency has theoretical interest and practical applications for lexicography and word sense disambiguation. Although WordNet and SemCor contain some information about sense frequency in English, it is not enough for either practical or research purposes. This information is even lacking in Russian. To fill this lacuna, we developed and tested an automated system based on semantic vectors, which deals with the problem of sense frequency for Russian nouns. The model is first trained unsupervised on large corpora and then supplied with contexts and collocations from the Active Dictionary of Russian. The dictionary examples are used either for supervised post-training or for automatic labeling of clusters that are learned unsupervised. This allows us to reach a frequency estimation error of 11-15 percent on different corpora without additional labeled data. Word sense frequency distributions for 440 nouns are available online.

 

 

Language: English
DOI
Keywords: word2vecword sense disambiguationsemantic vectorss​emantics​polysemyw​ord sense frequencyWSD

In book

Quantitative approaches to the Russian language
Quantitative approaches to the Russian language
Abingdon: Routledge, 2018.
Similar publications
Высокоуровневая семантическая интерпретация структуры статических моделей для русского языка
Serikov O., Ganeeva V., Аксенова А. А. et al., Вестник Новосибирского государственного университета. Серия: Лингвистика и межкультурная коммуникация 2023 Т. 21 № 1 С. 67–82
Since its inception, the Word2vec vector space has become a universal tool both for scientific and practical activities. Over time, it became clear that there is a lack of new methods for interpreting the location of words in vector spaces. The existing methods included consideration of analogies or clustering of a vector space. In recent ...
Added: April 28, 2025
You shall know a piece by the company it keeps. Chess plays as a data for word2vec models
Orekhov B., / Series Computer Science "arxiv.org". 2024.
In this paper, I apply linguistic methods of analysis to non-linguistic data, chess plays, metaphorically equating one with the other and seeking analogies. Chess game notations are also a kind of text, and one can consider the records of moves or positions of pieces as words and statements in a certain language. In this article ...
Added: August 8, 2024
Effectiveness of ELMo embeddings, and semantic models in predicting review helpfulness
Malik M. S., Nawaz A., Jamjoom M. M. et al., Intelligent Data Analysis 2024 Vol. 28 No. 4 P. 1045–1065
Online product reviews (OPR) are a commonly used medium for consumers to communicate their experiences with products during online shopping. Previous studies have investigated the helpfulness of OPRs using frequency-based, linguistic, meta-data, readability, and reviewer attributes. In this study, we explored the impact of robust contextual word embeddings, topic, and language models in predicting the ...
Added: February 26, 2024
Конструирование образа города в официальной и обыденной коммуникации: сравнительный анализ (на материале социальных медиа)
Matkin N., Коммуникации. Медиа. Дизайн 2025 Т. 10 № 3 С. 89–110
The article offers an analysis and visualization of Russian city images that emerge in the comments of urban community subscribers and posts from administrative press services. The city image is regarded as a frame structure that develops through political and interpersonal communication in the network. The social component of the city image is identified as ...
Added: November 15, 2023
Identifying emerging trends and hot topics through intelligent data mining: the case of clinical psychology and psychotherapy
Sokolova A., Lobanova P., Kuzminov I., Foresight 2024 Vol. 26 No. 1 P. 155–180
Purpose The purpose of the paper is to present an integrated methodology for identifying trends in a particular subject area based on a combination of advanced text mining and expert methods. The authors aim to test it in an area of clinical psychology and psychotherapy in 2010–2019. Design/methodology/approach The authors demonstrate the way of applying text-mining and the ...
Added: October 12, 2023
How to detect propaganda from social media? Exploitation of semantic and fine-tuned language models
Malik M. S., Imran T., Mona Mamdouh J., PeerJ Computer Science 2023 Vol. 9 Article e1248
Online propaganda is a mechanism to influence the opinions of social media users. It is a growing menace to public health, democratic institutions, and public society. The present study proposes a propaganda detection framework as a binary classification model based on a news repository. Several feature models are explored to develop a robust model such ...
Added: September 4, 2023
Automated defect identification for cell phones using language context, linguistic and smoke-word models
Muhammad Z. Y., Malik M. S., Ignatov D. I., Expert Systems with Applications 2023 Vol. 227 Article 120236
Product defects are a widespread concern for manufacturers when conducting quality and customer relationship management. Prior approaches addressed many electronic products however cell phones are still unexplored. Moreover, prior work mainly focused on the lexicon, probabilistic graphic, failure mode, and effect analysis models but the utilization of word embeddings and language models are not explored. State-of-the-art contextual word embeddings and language models generate automated features and ...
Added: June 13, 2023
Sense-Annotated Corpus for Russian
Kirillovich A., Loukachevitch N. V., Kulaev M. et al., , in: Proceedings of the Fifth International Conference Computational Linguistics in Bulgaria (CLIB 2022).: Sofia: Bulgarian Academy of Sciences, 2022. P. 130–136.
We present a sense-annotated corpus for Russian. The resource was obtained my manually annotating texts from the OpenCorpora corpus, an open corpus for the Russian language, by senses of Russian wordnet RuWordNet. The annotation was used as a test collection for comparing unsupervised (Personalized Pagerank) and pseudo-labeling methods for Russian word sense disambiguation. ...
Added: September 8, 2022
Detection of semantic changes in Russian nouns with distributional models and grammatical features
Ryzhova A., Ryzhova D., Sochenkov I., , in: Computational Linguistics and Intellectual Technologies: Papers from the Annual International Conference “Dialogue” (2021)Issue 20: Основной том.: -, 2021. P. 597–606.
Added: October 30, 2021
An Unsupervised Method for Weighting Finite-state Morphological Analyzers
Tyers F. M., Keleg A., Pirinen T., , in: Proceedings of The 12th Language Resources and Evaluation ConferenceVol. 12.: European Language Resources Association (ELRA), 2020. P. 3842–3850.
Morphological analysis is one of the tasks that have been studied for years. Different techniques have been used to develop models for performing morphological analysis. Models based on finite state transducers have proved to be more suitable for languages with low available resources. In this paper, we have developed a method for weighting a morphological ...
Added: April 20, 2021
Automated Analysis of Discourse Coherence in Schizophrenia: Approximation of Manual Measures
Ryazanskaya G., Khudyakova M., , in: Proceedings of the LREC 2020 Workshop on: Resources and Processing of Linguistic, Para-linguistic and Extra-linguistic Data from People with Various Forms of Cognitive/Psychiatric/Developmental Impairments (RaPID-3).: European Language Resources Association (ELRA), 2020. P. 98–107.
Disorganized, or incoherent, speech is one of the important criteria for diagnosing schizophrenia. However, there is still a lack of a rather quick objective method of measuring speech coherence. Automated discourse analysis is a possible solution to this problem. We analyzed discourse coherence in a set of spoken narratives by people with schizophrenia and neurotypical speakers ...
Added: February 2, 2021
Learning Word Embeddings without Context Vectors
Zobnin A., Elistratova E., , in: Proceedings of the 4th Workshop on Representation Learning for NLP (RepL4NLP-2019)Issue W19-43.: Association for Computational Linguistics, 2019. P. 244–249.
Most word embedding algorithms such as word2vec or fastText construct two sort of vectors: for words and for contexts. Naive use of vectors of only one sort leads to poor results. We suggest using indefinite inner product in skip-gram negative sampling algorithm. This allows us to use only one sort of vectors without loss of ...
Added: November 9, 2019
WORD VECTOR MODELS AS AN OBJECT OF LINGUISTIC RESEARCH
Shavrina T., , in: Компьютерная лингвистика и интеллектуальные технологии: По материалам ежегодной международной конференции «Диалог» (Москва, 29 мая — 1 июня 2019 г.)Вып. 18(25).: [б.и.], 2019. P. 576–588.
This article launches a series of studies in which popular vector word2vec models are considered not as an element of the architecture of an NLP application, but as an independent object of linguistic research. The linguist's view on the surrogate of contexts on the corpus, as which vector models can be considered, makes it possible ...
Added: September 5, 2019
Extraction of Hypernyms from Dictionaries with a Little Help from Word Embeddings
Karyaeva M., Braslavski P., Kiselev Y., , in: Analysis of Images, Social Networks and Texts. 7th International Conference AIST 2018.: Springer, 2018. P. 76–87.
The paper investigates several techniques for hypernymy extraction from a large collection of dictionary definitions in Russian. First, definitions from different dictionaries are clustered, then single words and multiwords are extracted as hypernym candidates. A classification-based approach on pre-trained word embeddings is implemented as a complementary technique. In total, we extracted about 40K unique hypernym ...
Added: March 11, 2019
Отрицательная и положительная поляризация: семантические источники
Apresyan V., В кн.: Компьютерная лингвистика и интеллектуальные технологии: По мате­риалам ежегодной международной конференции «Диалог» (Москва, 31 мая — 3 июня 2017 г.). Вып. 16 (23): В 2 т.Т. 2.: М.: Изд-во РГГУ, 2017. Гл. 1 С. 2–16.
Negative and positive polarity items (NPIs and PPIs) are one of the well-explored topics in formal semantics and typology. However, the phenomenon of polarization is only addressed on a very limited linguistic material, such as indefinite pronouns (some vs. any), temporal adverbs (yet vs. already), certain idioms (not to lift one’s finger), expressions of attitude ...
Added: August 29, 2018
Семантика качественных прилагательных в гойдельских языках: ‘тяжелый’ и ‘легкий’
Dereza O., В кн.: Сборник научно-исследовательских работ по итогам конкурса НИРС НИУ ВШЭ – 2015.: М.: Издательский дом НИУ ВШЭ, 2016. С. 210–225.
This is a small corpus study of Goidelic adjectives denoting physical qualities of heaviness and lightness, namely trom and éadrom in Irish and trom and aotrom (eutrom) in Scottish Gaelic, which both go back to Old Irish forms tromm and étromm. Obviously, étromm is derived from tromm with a negative prefix é, which suggests a ...
Added: October 5, 2017
Word Sense Frequency Estimation for Russian: Verbs, Adjectives, and Different Dictionaries
Lopukhina A., Лопухин К. А., , in: Electronic lexicography in the 21st century. Proceedings of eLex 2017 conference.: Brno: Lexical Computing CZ s.r.o., 2017. P. 267–280.
In this paper, we investigate several extensions to our prior work on sense frequency estimation for Russian. Our method is based on semantic vectors and is able to achieve good accuracy for sense frequency estimation trained on dictionary entries from the Active Dictionary of Russian and unannotated corpora. We apply our method to verbs and ...
Added: September 27, 2017
Word Sense Induction for Russian: Deep Study and Comparison with Dictionaries
Лопухин К. А., Iomdin B., Lopukhina A., Компьютерная лингвистика и интеллектуальные технологии 2017 Vol. 1 No. 16 P. 121–134
The assumption that senses are mutually disjoint and have clear boundaries has been drawn into doubt by several linguists and psychologists. The problem of word sense granularity is widely discussed both in lexicographic and in NLP studies. We aim to study word senses in the wild—in raw corpora— by performing word sense induction (WSI). WSI ...
Added: September 27, 2017
Webvectors: A toolkit for building web interfaces for vector semantic models
Kutuzov A., Kuzmenko E., , in: Supplementary Proceedings of the 5th International Conference on Analysis of Images, Social Networks and Texts (AIST-SUP 2016), Yekaterinburg, Russia, April 7-9, 2016.Vol. 1710.: Aachen: CEUR Workshop Proceedings, 2016. P. 155–161.
The paper presents a free and open source toolkit which aim is to quickly deploy web services handling distributed vector models of semantics. It fills in the gap between training such models (many tools are already available for this) and dissemination of the results to general public. Our toolkit, WebVectors, provides all the necessary routines for ...
Added: April 20, 2017
Improving Distributional Semantic Models Using Anaphora Resolution during Linguistic Preprocessing
Kutuzov A. B., Козлова О. С., , in: Компьютерная лингвистика и интеллектуальные технологии: По материалам ежегодной международной конференции «Диалог» (Москва,1–4 июля 2016 г.)Вып. 15.: М.: Изд-во РГГУ, 2016. P. 288–300.
In natural language processing, distributional semantic models are known as an efficient data driven approach to word and text representation, which allows computing meaning directly from large text corpora into word embeddings in a vector space. This paper addresses the role of linguistic preprocessing in enhancing performance of distributional models, and particularly studies pronominal anaphora ...
Added: November 12, 2016
  • About
  • About
  • Key Figures & Facts
  • Sustainability at HSE University
  • Faculties & Departments
  • International Partnerships
  • Faculty & Staff
  • HSE Buildings
  • HSE University for Persons with Disabilities
  • Public Enquiries
  • Studies
  • Admissions
  • Programme Catalogue
  • Undergraduate
  • Graduate
  • Exchange Programmes
  • Summer University
  • Summer Schools
  • Semester in Moscow
  • Business Internship
  • Research
  • International Laboratories
  • Research Centres
  • Research Projects
  • Monitoring Studies
  • Conferences & Seminars
  • Academic Jobs
  • Yasin (April) International Academic Conference on Economic and Social Development
  • Media & Resources
  • Publications by staff
  • HSE Journals
  • Publishing House
  • iq.hse.ru: commentary by HSE experts
  • Library
  • Economic & Social Data Archive
  • Video
  • HSE Repository of Socio-Economic Information
  • HSE1993–2026
  • Contacts
  • Copyright
  • Privacy Policy
  • Site Map
Edit