?
RusLan-M: Technical Design and Processing Pipeline of a Longitudinal Multimedia Corpus of Russian Child Speech
This paper presents RusLan-M (v.1.0), an open-access longitudinal multimedia corpus of spontaneous early child speech in Russian, designed for corpus-based research. The corpus consists of video recordings of naturalistic childcaregiver interaction, transcribed and annotated in the CHAT format and publicly available in the CHILDES database. We describe the process of corpus creation, data collection, cleaning and anonymization procedures, as well as transcription and annotation principles. The current version comprises 41 hours of recordings and over 35,000 child utterances, enabling the investigation of longitudinal trajectories of lexical growth, morphological development, syntactic complexity, and patterns of child-directed speech in Russian. Future development directions include corpus expansion with additional longitudinal datasets, systematic manual validation of automatic annotation tiers, and further integration of automatic alignment and annotation tools (BatchAlign2) adapted to Russian child speech. RusLan-M is designed for research on language acquisition, corpus linguistics, and the development of computational methods for morphologically rich languages.