?
The Emotional Gap: Is Sentiment Analysis Enough for Fiction? Evidence from the Corpus of Russian Short Stories of the Early 20th Century
The paper presents the results of a series of studies on the emotional dimension of the Russian short story of the first third of the 20th century. The study is based on a sample of 210 texts from the Corpus of the Russian Short Story of 1900–1930. Two fundamentally different approaches are compared: automatic detection of emotion-related patterns in literary texts using computational linguistics tools, and the measurement of readers’ emotional responses through expert annotations. In the first approach, three sentiment analysis methods are compared: lexicon-based, machine learning-based, and distributional semantics-based. The results show that correlations between these methods on literary texts are low, suggesting that each captures a different aspect of emotionality. In the second approach, patterns of emotional perception are analysed on the basis of annotation by contemporary readers with a background in philology. A comparison of lexical and perceptual emotionality assessment shows that reader response to different emotions is not equally tied to their presence in the vocabulary of the text: some emotions, such as sadness, are felt more strongly where the corresponding words are frequent, while others, such as happiness and sometimes surprise and anger, arise mainly from plot and context rather than from individual words. The emotionality of a literary text is thus a complex, multidimensional phenomenon that cannot be fully captured through vocabulary alone, and hybrid approaches combining computational and literary methods are needed for its study.