?
When Texts and Visuals Disagree: Multimodal Incongruence in Sarcasm
Detection of sarcasm is a complex task, and several state-of-the-art approaches have been proposed to address it. In the case of multimodal data, sarcasm can often be identified through incongruence between modalities. In this paper, we investigate polarity mismatch between visual and textual signals using data from two sarcasm detection datasets: MUStARD (Castro et al., 2019) and MMSD2.0 (Ying et al., 2023). In addition, we annotate the data according to a pragmatic taxonomy in order to complement purely computational detection with a theoretical framework and explore the relationship between multimodal incongruence and different pragmatic types of sarcasm. Our preliminary results suggest that multimodal mismatch is present in both datasets but does not function as a universal property of sarcastic discourse. The relationship between mismatch and pragmatic categories also appears to be domain-sensitive: the most notable tendency is observed for whimper in MMSD2.0, whereas no comparable effect emerges in MUStARD. By contrast, text-only sarcasm probability does not vary significantly across the annotated categories. These findings suggest that multimodal mismatch may capture aspects of sarcastic realization that are not reflected in text-only sarcasm scores.