?
Проблема идентификации именованных сущностей при их автоматическом извлечении
The article is devoted to the overview of the basic properties of Named Entities Recognition (NER) system based on users’ dictionaries. The NER module is used in many applications. One of the promising applications is the usage of NER systems in order to enhance structured Semantic Web data (for instance, Linked Open Data ontologies) with the information extracted from unstructured texts. The focus of the paper is the methods of ambiguity resolution based on dictionaries and heuristic rules. The dictionary-oriented approach is motivated by the set of strict initial requirements. Firstly, the target set of Named Entities should be extracted with very high precision. Secondly, the system should be easily adapted to a new domain by non-specialists. Thirdly, these updates should result in the same high precision. We focus on the architecture of the dictionaries and on the properties that the dictionaries should have for each class of Named Entities. This serves to resolve ambiguous situations. The properties and structure of synonyms and context words, expressions and entities necessary for disambiguation are discussed.