?
From Data to Signs: A Foundation Model for Multilingual Sign Language Recognition
Video-based Isolated Sign Language Recognition (ISLR) problem presents significant challenges in scaling across diverse languages due to data scarcity and the computational costs associated with training of language-specific models. In this paper, we introduce a novel training pipeline that leverages self-supervised learning on a large-scale sign language dataset. To obtain the foundation model, we utilize the VideoMAE architecture with a ViT-L backbone, pre-trained on the Kinetics-400 dataset. In particular, to capture the fine-grained spatialtemporal features essential for sign language processing, we adopted a tube masking mechanism, in which the input video is split into spatiotemporal tubes with 90% masking coefficient. The targeted fine-tuning of this model is implemented for easy adaptation to multiple sign languages with limited number of training videos. Experimental results demonstrate the benefits of our approach, achieving near-state-of-the-art results for Russian, American, Greek, and Turkish sign languages. Notably, we achieve high accuracy in up to 3.5 times less number of training epochs per language compared to conventional training from scratch, leading to significant reduction of training time and resource requirements, and, hence, facilitating development of high-performance ISLR models for various sign languages.