?
Automated Word Sense Frequency Estimation for Russian Nouns
According to G. K. Zipf’s observation, there is a strong correlation between word frequency and polysemy. Yet word sense frequency distribution is a neglected area in computational linguistics. Furthermore, the study of sense frequency has theoretical interest and practical applications for lexicography and word sense disambiguation. Although WordNet and SemCor contain some information about sense frequency in English, it is not enough for either practical or research purposes. This information is even lacking in Russian. To fill this lacuna, we developed and tested an automated system based on semantic vectors, which deals with the problem of sense frequency for Russian nouns. The model is first trained unsupervised on large corpora and then supplied with contexts and collocations from the Active Dictionary of Russian. The dictionary examples are used either for supervised post-training or for automatic labeling of clusters that are learned unsupervised. This allows us to reach a frequency estimation error of 11-15 percent on different corpora without additional labeled data. Word sense frequency distributions for 440 nouns are available online.