?
Entity Linking over Nested Named Entities for Russian
P. 4458–4466.
In this paper, we describe entity linking annotation over nested named entities in the recently released Russian NEREL dataset for information extraction. The NEREL collection is currently the largest Russian dataset annotated with entities and relations. It includes 933 news texts with annotation of 29 entity types and 49 relation types. The paper describes the main design principles behind NEREL’s entity linking annotation, provides its statistics, and reports evaluation results for several entity linking baselines. To date, 38,152 entity mentions in 933 documents are linked to Wikidata. The NEREL dataset is publicly available.
Language:
English
Keywords: information extraction
Publication based on the results of:
In book
Marseille: European Language Resources Association (ELRA), 2022.
F. M. Grozovskiy, I. V. Loginova, Automatic Documentation and Mathematical Linguistics 2025 Vol. 59 No. 4 P. 269–278
The paper proposes an approach to the automated extraction and structuring of information from
text, combining web scraping for data collection from online sources with a large language model for subsequent
data mining. As a case study, texts from news publications on technology readiness levels from the
CNews website were chosen to test the developed methodology in a ...
Added: August 25, 2025
Baimuratov I., Karpovich A., Lisanyuk E. et al., , in: JCDL '24: Proceedings of the 24th ACM/IEEE Joint Conference on Digital Libraries.: NY: Association for Computing Machinery (ACM), 2024. Ch. 6.
Peer review is a cornerstone of the academic editorial decisionmaking process, yet it faces significant challenges. Artificial intelligence can help address these challenges, but its use raises concerns about reliability and the potential for reproducing existing biases. In this research, we employ a formal argumentation-theoretic framework that allows for explicit analysis of arguments and their ...
Added: May 29, 2025
Bolshakova E. I., Семак В. В., Интеллектуальные системы. Теория и приложения 2021 Т. 25 № 4 С. 239–242
An approach to automatic extraction of terms from an individual scientific text is reported, which combines known methods: linguistic patterns, statistical terminological measures, methods of graph ranking. The combined methods and stages for extracting, selection and ranking of terms are described, which are implemented for processing documents in Russian. The results of experiments on extracting ...
Added: November 23, 2023
Loukachevitch N., Artemova E., Batura T. et al., Language Resources and Evaluation 2024 Vol. 58 P. 547–583
This paper describes NEREL—a Russian news dataset suited for three tasks: nested named entity recognition, relation extraction, and entity linking. Compared to flat entities, nested named entities provide a richer and more complete annotation while also increasing the coverage of relations annotation and entity linking. Relations between nested named entities may cross entity boundaries to ...
Added: September 24, 2023
Tutubalina E., Алимова И. С., Мифтахутдинов З. et al., Bioinformatics 2021 Vol. 37 No. 2 P. 243–249
Drugs and diseases play a central role in many areas of biomedical research and healthcare. Aggregating knowledge about these entities across a broader range of domains and languages is critical for information extraction (IE) applications. To facilitate text mining methods for analysis and comparison of patient’s health conditions and adverse drug reactions reported on the ...
Added: January 13, 2021
Chernyavskiy A., Ilvovsky D., , in: Proceedings of the Second Workshop on Fact Extraction and VERification (FEVER).: Association for Computational Linguistics, 2019. P. 69–78.
Triggered by Internet development, a large amount of information is published in online sources. However, it is a well-known fact that publications are inundated with inaccurate data. That is why fact-checking has become a significant topic in the last 5 years. It is widely accepted that factual data verification is a challenge even for the ...
Added: September 15, 2020
Alsu Zaynutdinova, Dina Pisarevskaya, Zubov M. et al., , in: Proceedings of the Fifth Workshop on Experimental Economics and Machine Learning at the National Research University Higher School of Economics co-located with the Seventh International Conference on Applied Research in Economics (iCare7).: Aachen: CEUR Workshop Proceedings, 2019. P. 121–127.
Russian Federation and European Union are fighting against fake news together with other countries in various topics. The disinformation affected British referendum of existing EU, the US election and Catalonia’s referendum are broadly studied. A need for automated factchecking increases, European Commission’s Action Plan 8 is an evidence. In this work, we develop a model ...
Added: November 19, 2019
Wohlgenannt G., von Waldenfels R., Toldova S. et al., Manchester: EasyChair, 2019.
The EPiC Series in Language and Linguistics publishes high quality collections of papers in language, linguistics and related areas. ...
Added: September 9, 2019
Azerkovich I., , in: Artificial Intelligence and Natural Language, 7th International Conference, AINL 2018, St. Petersburg, Russia, October 17–19, 2018, ProceedingsIssue 930.: Switzerland: Springer, 2018. P. 107–112.
Semantic information has been deemed a valuable resource for
solving the task of coreference resolution by many researchers. Unfortunately,
not much has been done in the direction of using this data when working with
Russian data. This work describes the first step of a research, attempting to create
a coreference resolution system for Russian based on semantic data, concerned
with ...
Added: September 5, 2018
Starostin A. S., Bocharov V. V., Alexeeva S. V. et al., , in: Компьютерная лингвистика и интеллектуальные технологии: По материалам ежегодной международной конференции «Диалог» (Москва,1–4 июля 2016 г.)Вып. 15.: М.: Изд-во РГГУ, 2016. P. 688–705.
In this paper, we describe the rules and results of the FactRuEval informa- tion extraction competition held in 2016 as part of the Dialogue Evaluation initiative in the run-up to Dialogue 2016. The systems were to extract in- formation from Russian texts and competed in two named entity extraction tracks and one fact extraction track. ...
Added: October 7, 2016
Skorinkin D.A., Budnikov E. A., Stepanova M. E. et al., Компьютерная лингвистика и интеллектуальные технологии 2016 No. 15 P. 721–733
This paper presents a rule-based approach to Information Extraction (IE) task within FactRuEval-2016 competition. Our system is based on ABBYY Compreno Technology. The technology uses the results of deep syntactic-semantic analysis, which leads to significant reduction of the number of necessary rules and makes them laconic. The evaluation was conducted on FactRuEval dataset. FactRuEval is ...
Added: August 28, 2016
Switzerland: Springer, 2014.
This book constitutes the refereed proceedings of the 6th IAPR TC3 International Workshop on Artificial Neural Networks in Pattern Recognition, ANNPR 2014, held in Montreal, QC, Canada, in October 2014. The 24 revised full papers presented were carefully reviewed and selected from 37 submissions for inclusion in this volume. They cover a large range of ...
Added: September 30, 2014
Соколова Е. Г., Сосенская Т. Б., Toldova S. et al., В кн.: Типология и теория языка: от описания к объяснению.: М.: Языки русской культуры, 1999. С. 600–612.
Статья посвящена системе автоматического извлечения информации из текста. ...
Added: April 7, 2014