?
Faroese Corpus Development: Strategies for Low-Resource Languages Corpora
This paper presents the development of a comprehensive morphologically annotated corpus for Faroese, a low-
resource Germanic language. We describe the creation of a large corpus of contemporary news texts automatically
annotated using a custom-trained SpaCy model. The study demonstrates the effectiveness of creating linguistic re -
sources for low-resource languages using minimal initial data. We trained a Transformer-based morphological pars-
ing model on the small but high-quality OFT treebank using 5-fold cross-validation, achieving significant accuracy
in morphological tagging and lemmatization. Manual evaluation confirms satisfactory performance of the automatic
annotation, though certain challenges remain in distinguishing homonymous word forms across different parts of
speech. This research provides a methodological framework for developing comprehensive linguistic resources for
other low-resource languages with minimal initial data requirements