?
RusLan-M: longitudinal multimedia corpus of early child speech in Russian
This article presents the Russian Language-Monolingual corpus (RusLan-M, v.1.0), a longitudinal multimedia collection of early child speech from two Russian-speaking monolingual children: Tosya (ages 0;10–3;10, 246 recordings) and Yasha (ages 1;04–3;00, 42 recordings). The corpus consists of approximately 41 h (2,454 min.) of video recordings and 35,386 child utterances, available with transcriptions in the CHAT format on TalkBank. The corpus adheres to strict ethics requirements for data sharing with anonymization. We also conducted two exploratory investigations of the acquisition of Russian morphology using the mean length of utterance (MLU) and the newly developed Index of Productive Syntax for evaluating grammatical complexity in the nominal system of Russian (IPSyn-NP-R). These investigations illustrate how a comprehensive analysis of syntactic and morphological structures in Russian language development can be conducted and highlight the kinds of questions that can be answered using the RusLan-M data. The RusLan-M corpus represents an application of corpus linguistics methods to the Russian language and fills a significant gap in the limited data and resources available for Russian child language research.