?
An Audio-Gated Cascade for Resource-Efficient B imodalEmotionRecognition
Bimodal (audio–visual) emotion recognition systems typ-ically process speech and video in parallel with heavy neural back-bones, incurring a computational cost that makes them unavailable on mobile, embedded and edge platforms. In this paper we propose a two-stage cascade architecture that decouples the two modalities in time. A lightweight binary audio gate, implemented as a compact log-mel CNN, first decides whether an utterance is neutral; only when emo-tional speech is detected does a resource-intensive bimodal head clas-sify the specific emotion. The same log-mel spectrogram is reused by both the gate and the audio branch of the bimodal head, removing the need for large self-supervised audio encoders and shrinking the model to approximately 37.3 MB – almost ten times smaller than the clos-est bimodal competitor.