• A
  • A
  • A
  • АБВ
  • АБВ
  • АБВ
  • A
  • A
  • A
  • A
  • A
Обычная версия сайта
  • RU
  • EN
  • HSE University
  • Publications
  • Articles
  • A Bimodal Approach for Speech Emotion Recognition using Audio and Text
  • RU
  • EN
Расширенный поиск
Высшая школа экономики
Национальный исследовательский университет
Priority areas
  • business informatics
  • economics
  • engineering science
  • humanitarian
  • IT and mathematics
  • law
  • management
  • mathematics
  • sociology
  • state and public administration
by year
  • 2028
  • 2027
  • 2026
  • 2025
  • 2024
  • 2023
  • 2022
  • 2021
  • 2020
  • 2019
  • 2018
  • 2017
  • 2016
  • 2015
  • 2014
  • 2013
  • 2012
  • 2011
  • 2010
  • 2009
  • 2008
  • 2007
  • 2006
  • 2005
  • 2004
  • 2003
  • 2002
  • 2001
  • 2000
  • 1999
  • 1998
  • 1997
  • 1996
  • 1995
  • 1994
  • 1993
  • 1992
  • 1991
  • 1990
  • 1989
  • 1988
  • 1987
  • 1986
  • 1985
  • 1984
  • 1983
  • 1982
  • 1981
  • 1980
  • 1979
  • 1978
  • 1977
  • 1976
  • 1975
  • 1974
  • 1973
  • 1972
  • 1971
  • 1970
  • 1969
  • 1968
  • 1967
  • 1966
  • 1965
  • 1964
  • 1963
  • 1958
  • More
Subject
News
September 11, 2026
How to Assess Students Knowledge in the Age of AI
A researcher at HSE University has proposed a flowchart to help lecturers decide how to assess students who use artificial intelligence. It shows where the use of AI should be restricted and where it can be incorporated into the learning process. The article has been published in IT Professional.
September 9, 2026
‘Balkan Hospitality Opens Doors: Studying Dialects on the Verge of Extinction
You cannot study spoken dialects from books. Instead, you need to go to a village, seek out its elders, and earn the trust of local residents before you can record hours of spontaneous stories. This is how Natalia Muravleva, Associate Professor at the Faculty of Humanities, conducts her research. Her internship in Serbia continued her long-standing study of dialects spoken by Macedonian settlers. In this interview, she discusses how diaspora cultural centres help researchers reach informants, why native speakers need to be interviewed only in their own language (otherwise, as she puts it, they may 'break'), and how a single field season helped her finalise her monograph. She also shares warm memories of autumn in Belgrade and of colleagues with whom grammar can be discussed in three languages at once.
September 9, 2026
Scientists Train Neural Network to Generate Process Plans from 3D Models
Researchers at the HSE FCS AI and Digital Science Institute have developed CAD2TechSpec, a framework that converts 3D models of mechanical parts into machining process plans—step-by-step instructions for machine tools. The solution aims to reduce the time required for the design and preparation of technical process documentation in mechanical engineering, aircraft manufacturing, and other high-tech industries. The study findings have been published in PeerJ Computer Science.

 

Have you spotted a typo?
Highlight it, click Ctrl+Enter and send us a message. Thank you for your help!

Publications
  • Books
  • Articles
  • Chapters of books
  • Working papers
  • Report a publication
  • Research at HSE

?

A Bimodal Approach for Speech Emotion Recognition using Audio and Text

Journal of Internet Services and Information Security. 2021. No. 1. P. 80–96.
Verkholyak O., Dvoynikova A., Karpov A.

This paper presents a novel bimodal speech emotion recognition system based on analysis of acoustic and linguistic information. We propose a novel decision-level fusion strategy that leverages both emotions and sentiments extracted from audio and text transcriptions of extemporaneous speech utterances. We perform experimental study to prove the effectiveness of the proposed methods using emotional speech database RAMAS, revealing classification results of 7 emotional states (happy, surprised, angry, sad, scared, disgusted, neutral) and 3 sentiment categories (positive, negative, neutral). We compare relative performance of unimodal vs. bimodal systems, analyze their effectiveness on different levels of annotation agreement, and discuss the effect of reduction of training data size on the overall performance of the systems. We also provide important insights about contribution of each modality for the best optimal performance for emotions classification, which reaches UAR=72.01% on the highest 5-th level of annotation agreement.

Language: English
Full text
DOI
Text on another site
Keywords: Sentiment analysisspeech emotion recognitioncomputational paralinguisticsAnnotation agreement Bimodal fusion
Similar publications
Enhancing Emotion Recognition in Speech Based on Self-Supervised Learning: Cross-Attention Fusion of Acoustic and Semantic Features
Deeb B., Andrey V. Savchenko, Makarov I., IEEE Access 2026 Vol. 13 P. 56283–56295
Speech Emotion Recognition has gained considerable attention in speech processing and machine learning due to its potential applications in human-computer interaction, mental health monitoring, and customer service. However, state-of-the-art models for speech emotion recognition use many parameters, which leads to computational complexity. In this paper, we introduce a novel deep-learning model to enhance the accuracy ...
Added: June 16, 2026
Emotion Recognition and Sentiment Analysis of Extemporaneous Speech Transcriptions in Russian
Dvoynikova A., Verkholyak O., Karpov A., Lecture Notes in Computer Science 2020 Vol. 12335 LNAI P. 136–144
Speech can be characterized by acoustical properties and semantic meaning, represented as textual speech transcriptions. Apart from the meaning content, textual information carries a substantial amount of paralinguistic information that makes it possible to detect speaker’s emotions and sentiments by means of speech transcription analysis. In this paper, we present experimental framework and results for ...
Added: April 24, 2026
A framework for text mining on Twitter: a case study on joint comprehensive plan of action (JCPOA)- between 2015 and 2019
Behzadidoost R., Quality and Quantity 2021 Vol. 56 No. 5 P. 3053–3084
In the big data era, there is a necessity for effective frameworks to collect, retrieve, and manage data. As not all tweets are hashtagged by users, retrieving them is a complicated task. To address this issue, we present a rule-based expert system classifier that uses the well-known concept of fingerprint in the judicial sciences. This ...
Added: March 27, 2026
Analyzing and forecasting P/E ratios using investor sentiment in panel data regression and LSTM models
Dolaeva A., Beliaeva U., Dmitry Grigoriev et al., International Review of Economics and Finance 2025 Vol. 98 Article 103840
Added: July 11, 2025
CA-SER: Cross-Attention Feature Fusion for Speech Emotion Recognition
Deeb B., Savchenko A., Makarov I., , in: ECAI 2024. 27th European Conference on Artificial Intelligence, October 19 – 24 October 2024, Santiago de Compostela, Spain – Including 13th Conference on Prestigious Applications of Intelligent Systems (PAIS 2024).: IOS Press, 2024. P. 4479–4482.
In this paper, we introduce a novel tool for speech emotion recognition, CA-SER, that borrows self-supervised learning to extract semantic speech representations from a pre-trained wav2vec 2.0 model and combine them with spectral audio features to improve speech emotion recognition. Our approach involves a self-attention encoder on MFCC features to capture meaningful patterns in audio ...
Added: February 15, 2025
Alternative method sentiment analysis using emojis and emoticons
Surikov A., Evgeniia Egorova, Procedia Computer Science 2020 Vol. 178 P. 182–193
Our research aims to develop an alternative method for analyzing the tonality of the texts. Most of the traditional methods for determining tonality classes are based on text analysis and ignore various emotional indicators that users actively used in social networks. Therefore, it improves the quality of predicting the tonality of the class. The study ...
Added: May 15, 2024
The voice of Twitter: observable subjective well-being inferred from tweets in Russian
Smetanin S., Mikhail Komarov, PeerJ Computer Science 2022 Vol. 8 Article e1181
As one of the major platforms of communication, social networks have become a valuable source of opinions and emotions. Considering that sharing of emotions offline and online is quite similar, historical posts from social networks seem to be a valuable source of data for measuring observable subjective well-being (OSWB). In this study, we calculated OSWB ...
Added: December 29, 2022
Research Anthology on Implementing Sentiment Analysis Across Multiple Disciplines
IGI Global, 2022.
The rise of internet and social media usage in the past couple of decades has presented a very useful tool for many different industries and fields to utilize. With much of the world’s population writing their opinions on various products and services in public online forums, industries can collect this data through various computational tools ...
Added: August 3, 2022
Speaker-Aware Training of Speech Emotion Classifier with Speaker Recognition
Savchenko L., Savchenko A., , in: Speech and Computer. 23rd International Conference, SPECOM 2021, St. Petersburg, Russia, September 27–30, 2021Vol. 12997.: St. Petersburg: Springer, 2021. Ch. 55 P. 614–625.
Added: September 24, 2021
Using Intelligent Text Analysis of Online Reviews to Determine the Main Factors of Restaurant Value Propositions
Fainshtein E., Serova E., , in: Handbook of Research on Applied Data Science and Artificial Intelligence in Business and Industry.: IGI Global, 2021. Ch. 10 P. 223–240.
Added: July 24, 2021
On the Impact of Word Error Rate on Acoustic-Linguistic Speech Emotion Recognition: An Update for the Deep Learning Era
Sokolov A., / Series Computer Science "arxiv.org". 2021.
Text encodings from automatic speech recognition (ASR) transcripts and audio representations have shown promise in speech emotion recognition (SER) ever since. Yet, it is challenging to explain the effect of each information stream on the SER systems. Further, more clarification is required for analysing the impact of ASR's word error rate (WER) on linguistic emotion ...
Added: November 17, 2020
Social Network Sentiment Analysis and Message Clustering
Kharlamov A. A., Orekhov A., Bodrunova S. et al., Lecture Notes in Computer Science 2019 Vol. 11938 P. 18–31
Till today, classification of documents into negative, neutral, or positive remains a key task within the analysis of text tonality/sentiment. There are several methods for the automatic analysis of text sentiment. The method based on network models, the most linguistically sound, to our viewpoint, allows us take into account the syntagmatic connections of words. Also, ...
Added: October 29, 2020
  • About
  • About
  • Key Figures & Facts
  • Sustainability at HSE University
  • Faculties & Departments
  • International Partnerships
  • Faculty & Staff
  • HSE Buildings
  • HSE University for Persons with Disabilities
  • Public Enquiries
  • Studies
  • Admissions
  • Programme Catalogue
  • Undergraduate
  • Graduate
  • Exchange Programmes
  • Summer University
  • Summer Schools
  • Semester in Moscow
  • Business Internship
  • Research
  • International Laboratories
  • Research Centres
  • Research Projects
  • Monitoring Studies
  • Conferences & Seminars
  • Academic Jobs
  • Yasin (April) International Academic Conference on Economic and Social Development
  • Media & Resources
  • Publications by staff
  • HSE Journals
  • Publishing House
  • iq.hse.ru: commentary by HSE experts
  • Library
  • Economic & Social Data Archive
  • Video
  • HSE Repository of Socio-Economic Information
  • HSE1993–2026
  • Contacts
  • Copyright
  • Privacy Policy
  • Site Map
Edit