?
An experimental rule-based parser for Russian employing the NLP resources of the ETAP system
.
Inshakova E.S., Sizov V. G.
In book
Issue 19 (26). , ., 2020.
Shavrina T., Fenogenova A., Emelyanov A. et al., , in: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP).: Association for Computational Linguistics, 2020. P. 4717–4726.
In this paper, we introduce an advanced Russian general language understanding evaluation benchmark – RussianSuperGLUE. Recent advances in the field of universal language models and transformers require the development of a methodology for their broad diagnostics and testing for general intellectual skills - detection of natural language inference, commonsense reasoning, ability to perform simple logical ...
Added: June 14, 2026
Zmitrovich D., Abramov A., Kalmykov A. et al., , in: Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024).: ELRA and ICCL, 2024.
Transformer language models (LMs) are fundamental to NLP research methodologies and applications in various languages. However, developing such models specifically for the Russian language has received little attention. This paper introduces a collection of 13 Russian Transformer LMs, which spans encoder (ruBERT, ruRoBERTa, ruELECTRA), decoder (ruGPT-3), and encoder-decoder (ruT5, FRED-T5) architectures. We provide a report ...
Added: June 14, 2026
Biryukova K., Chelnokova D., Erkenova J. et al., , in: Analysis of Images, Social Networks and Texts. AIST 2024Issue 2364.: Cham: Springer, 2024. P. 109–121.
Visual Question Answering is one of the essential parts of machine reasoning. Datasets are created to train a model to perform this task. However, there are only a few datasets for the Russian language. Moreover, existing sets may have strong biases, allowing models to score high without reasoning. In this paper, we adapt the idea ...
Added: June 14, 2026
Chervyakov A., Isaeva U., Emelyanov A. et al., , in: Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers)Vol. 1.: Association for Computational Linguistics, 2026. P. 2114–2161.
Multimodal large language models (MLLMs) are currently at the center of research attention, showing rapid progress in scale and capabilities, yet their intelligence, limitations, and risks remain insufficiently understood. To address these issues, particularly in the context of the Russian language, where no multimodal benchmarks currently exist, we introduce MERA Multi, an open multimodal evaluation ...
Added: June 14, 2026
Association for Computational Linguistics, 2026.
Added: June 14, 2026
Chernogorskii F., Averkiev S., Kudraleeva L. et al., , in: Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 4: Student Research Workshop)Vol. 4.: Association for Computational Linguistics, 2026. P. 622–638.
This paper introduces DRAGOn, method to design a RAG benchmark on a regularly updated corpus. It features recent reference datasets, a question generation framework, an automatic evaluation pipeline, and a public leaderboard. Specified reference datasets allow for uniform comparison of RAG systems, while newly generated dataset versions mitigate data leakage and ensure that all models ...
Added: June 13, 2026
Kenneth E., Chung I., Kerboua I. et al., , in: Proceedings of the 13th International Conference on Learning Representations (ICLR 2025).: ICLR, 2025. P. 102004–102060.
Text embeddings are typically evaluated on a limited set of tasks, which are constrained by language, domain, and task diversity. To address these limitations and provide a more comprehensive evaluation, we introduce the Massive Multilingual Text Embedding Benchmark (MMTEB) - a large-scale, community-driven expansion of MTEB, covering over 500 quality-controlled evaluation tasks across 250+ languages. ...
Added: June 11, 2026
Снегирев А., Tikhonova M., Maksimova A. et al., , in: Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language TechnologiesVol. 1: Volume 1: Long Papers.: Association for Computational Linguistics, 2025. P. 236–254.
Embedding models play a crucial role in Natural Language Processing (NLP) by creating text embeddings used in various tasks such as information retrieval and assessing semantic text similarity. This paper focuses on research related to embedding models in the Russian language. It introduces a new Russian-focused embedding model called ru-en-RoSBERTa and the ruMTEB benchmark, the ...
Added: June 11, 2026
Churin I., Apishev M., Tikhonova M. et al., , in: Proceedings of the 6th Workshop on Computational Approaches to Discourse, Context and Document-Level Inferences (CODI 2025).: Suzhou: Association for Computational Linguistics, 2025. P. 1–13.
Recent progress in Natural Language Processing (NLP) has driven the creation of Large Language Models (LLMs) capable of tackling a vast range of tasks. A critical property of these models is their ability to handle large documents and process long token sequences, which has fostered the need for a robust evaluation methodology for long-text scenarios. ...
Added: June 11, 2026
Strube M., Braud C., Hardmeier C. et al., Suzhou: Association for Computational Linguistics, 2025.
Added: June 11, 2026
Rabat: Association for Computational Linguistics, 2026.
Added: May 19, 2026
Yalcin H., Demirhan D., Aracioglu B. et al., Technology in Society 2026 Vol. 84 Article 103094
This article comprehensively evaluates the critical role of FinTech in promoting carbon neutrality and green logistics practices in global supply chains. In our study, using bibliometric analysis, social network analysis and natural language processing (NLP) methods, we evaluate the potential of FinTech innovations to increase traceability, transparency and efficiency in supply chain processes. In this ...
Added: March 11, 2026
Association for Computational Linguistics, 2025.
The book contains this year’s edition of the Conference on Empirical Methods in Natural
Language Processing! Importantly, it marks the 30th edition of EMNLP. With over 8,000 submissions,
more than 3,000 accepted papers, and thousands of attendees, we have come a long way from that first
workshop, which had 14 accepted papers. As the field looks ahead, Suzhou ...
Added: November 16, 2025
Bolshakova E. I., Семак В. В., Программные продукты и системы 2025 Т. 38 № 1 С. 5–16
The current state in the field of automatic term extraction from specialized natural language texts, including scientific and technical documents, is considered. Practical applications of methods and tools for extracting terms from texts include creation of terminological dictionaries, thesauri, and glossaries of problem oriented domains, as well as extraction of keywords and construction of subject ...
Added: July 2, 2025
Springer, 2024.
This book constitutes the refereed proceedings of the 12th International Conference on Analysis of Images, Social Networks and Texts, AIST 2024, held in Bishkek, Kyrgyzstan, during October 17–19, 2024.
The 16 full papers included in this book were carefully reviewed and selected from 70 submissions. They were organized in topical sections as follows: Natural Language Processing; Computer Vision; Data Analysis and Machine Learning; ...
Added: May 29, 2025
Rome: Springer, 2025.
This book constitutes the refereed proceedings of the 15th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management, IC3K 2023, held in Rome, Italy, during November 13-15, 2023.
The 9 full papers and 8 short papers included in this book were carefully reviewed and selected from 166 submissions. They were organized in topical sections ...
Added: May 2, 2025
Association for Computational Linguistics, 2024.
CoNLL is a conference organized yearly by SIGNLL (ACL’s Special Interest Group on Natural Language Learning), focusing on theoretically, cognitively and scientifically motivated approaches to computational linguistics. This year, CoNLL was held alongside EMNLP 2024. ...
Added: March 11, 2025
Aleksandr Belov, Zakharov F., Litvinenko E. et al., , in: International IoT, Electronics and Mechatronics Conference, Volume 2. Proceedings of IEMTRONICS 2024. LNEE, volume 1228Vol. 1228.: Springer Publishing Company, 2025. P. 275–287.
Added: January 26, 2025
Malik M. S., Lecture Notes in Computer Science 2024 Vol. 14486 P. 3–17
In recent decades, hate speech on social media platforms has been on
the rise. It is highly desired to control this kind of material because it initiates
unrest and harms to the society. Literature describes several forms of the hate
speech and it is quite challenging to differentiate between these forms and to
design an automated detection system, especially ...
Added: December 12, 2024
Parakal E. G., Dudyrev E., Sergei O. Kuznetsov et al., Lecture Notes in Computer Science 2024 Vol. 14914 P. 286–301
This paper proposes an approach and an associated system
based on pattern structures, aimed at the classification of documents
represented as graphs. The representation of documents relies on Abstract
Meaning Representation (AMR) document graphs. Given a set of AMR
document graphs, the system learns characteristic graph patterns, that
can be reused by an aggregate rule classifier to predict the class ...
Added: September 10, 2024
Tyers Francis M., Gökırmak M., , in: Proceedings of the Fourth International Conference on Dependency Linguistics (DepLing, 2017).: Association for Computational Linguistics, 2017. P. 64–73.
This paper describes the development of the first syntactically annotated corpus of Kurmanji Kurdish. The corpus was used as one of the surprise languages in the 2017 CoNLL shared task on parsing Universal Dependencies. In the paper we describe how the corpus was prepared, some Kurmanji specific constructions that required special treatment, and we give ...
Added: September 30, 2017
Gareyshina A., Ionov M., Lyashevskaya O. et al., , in: Proceedings of COLING 2012: Posters.: Mumbai: The COLING 2012 Organizing Committee, 2012. P. 349–360.
The paper reports on the recent forum RU-EVAL - a new initiative for evaluation of Russian NLP resources, methods and toolkits. It started in 2010 with evaluation of morphological parsers, and the second event RU-EVAL 2012 (2011-2012) focused on syntactic parsing. Eight participating IT companies and academic institutions submitted their results for corpus parsing. We ...
Added: September 23, 2013