?
Некоторые инвариантные характеристики русской разговорной речи: фонетика, морфология, синтаксис (Some invariant featu res of Russian everyday speech: Phonology, morphology, syntax)
The presented research was carried out on the material of the ORD speech corpus in the framework of the project, dedicated to study sociolinguistic variation of Russian speech and aimed at identifying diagnostic features characterizing everyday speech of major social groups (age-, gender-, status-, professional-related, etc.). The obtained results showed that practically on each linguistic level one may observe the features exhibiting a very high similarity between different sociolects. In particular, the coincidence is observed in the distribution of phonemes, distribution of parts of speech, and the frequency of some syntactic structures. The distribution of phonemes was determined on the subcorpus of 172,000 allophones. The following ten phonemes are the most frequent in speech of all social groups: /a/ (18,18%), /i/ (9,04%), /t/ (6,36%), /o/ (5,43%), /u/ (4,49%), /n/ (4,11%), /j/ (3,82%), /e/ (3,57%), /k/ (3,35%), /s/ (3.01%). The distribution of parts of speech in everyday speech was obtained on the linguistically annotated subcorpus of 125,437 tokens and has the following breakdown: V (17,43%), S (15,29%), S-PRO (14,13%), PART (13,35%), CONJ (9,47%), PR (7,09%), ADV-PRO (5,30%), ADV (4,51%), A-PRO (4,30%), A (3,73%), PRAEDIC (1,84%), INTJ (1,41%), NUM (1,29%), PARENTH (0,56%), ANUM (0,27%), PRAEDIC-PRO (0,01%). At the syntactic level, one-element structures are prevailing in everyday speech of all social groups, the most frequent among them being D (particle / discursive word) (3,73%), S (2,26%), and V (1,88%). Statistical analysis of the left-branching and right-branching verb groups has showed that the first ones significantly prevail in speech of all sociolects. The revealed features reflect some constant, universal properties of everyday spoken Russian and can be used for adjustment and improvement of speech synthesis and recognition systems.