?
CAT&kittens: a corpus-based text-analytic tool for Russian academic writing
The development of the CAT corpus follows established corpus development procedures (e.g., BAWE). It was collected by extracting recently published texts sourced from textbooks, academic journals, and collecting high-quality master’s theses from available sources. All texts entered in CAT are divided into six disciplinary fields: social studies and history, political science and international relations, law, general and applied linguistics, economics, psychology and education science. Every discipline sub-corpus consists of about 300 to 400 thousand tokens, amounting to appr. 2 million tokens in the corpus in general. CAT is supplied with metalinguistic information, as well as morphological and syntactic annotation, carried out with the help of the annotation software RU Syntax (Mediankin et al. 2016). Further corpus improvement is also planned. Since the main goal of the project is to create a tool that compares novice texts to standard academic texts along the lists of pre-set criteria, the tool will run a series of “error analysis” test. The patterns of deviations are identified along lexical, collocational, morphological, and syntactic planes.