?
Cipher, transform, get lost: an anti-transparent system for distance measurement in East Slavic lects
Recent advances in computational historical linguistics have inspired a discussion on newly implemented quantitative methods — mainly, it is about their lack of transparency, and the ways to overcome it. This paper aims to demonstrate the advantages of transparency for such tools. The study compares two types of language distance measurement systems used in classification. Black-box systems transform the input data (such as the Swadesh list) into output data (language distance) with human- and machine-unexplainable decision-making. Language-agnostic systems (such as string similarity measures) analyse the input data and produce output data transparently, but do not consider the specifics of each language. For a proper comparison, I propose a new anti-transparent system based on hashing algorithms, vectorisation and language contact emulation. For my purposes, I use material from two test groups — East Slavic and Taa, both lexical and grammatical. East Slavic data are extracted from the corpora of Belogornoje, Megra, and Khislavichi and feature lists of Mokshenskaja, Kritskovschina and Pestschanka. Taa material consists of previously published Swadesh lists for the closely related !Xóõ (!Xoong), Kakia (Masarwa) and Nǀuǁen. An important new contribution of this work is the publication of new Swadesh wordlists for three East Slavic dialects.