• A
  • A
  • A
  • АБВ
  • АБВ
  • АБВ
  • A
  • A
  • A
  • A
  • A
Обычная версия сайта
  • RU
  • EN
  • HSE University
  • Publications
  • Book chapter
  • Parse thicket representations of text paragraphs
  • RU
  • EN
Расширенный поиск
Высшая школа экономики
Национальный исследовательский университет
Priority areas
  • business informatics
  • economics
  • engineering science
  • humanitarian
  • IT and mathematics
  • law
  • management
  • mathematics
  • sociology
  • state and public administration
by year
  • 2028
  • 2027
  • 2026
  • 2025
  • 2024
  • 2023
  • 2022
  • 2021
  • 2020
  • 2019
  • 2018
  • 2017
  • 2016
  • 2015
  • 2014
  • 2013
  • 2012
  • 2011
  • 2010
  • 2009
  • 2008
  • 2007
  • 2006
  • 2005
  • 2004
  • 2003
  • 2002
  • 2001
  • 2000
  • 1999
  • 1998
  • 1997
  • 1996
  • 1995
  • 1994
  • 1993
  • 1992
  • 1991
  • 1990
  • 1989
  • 1988
  • 1987
  • 1986
  • 1985
  • 1984
  • 1983
  • 1982
  • 1981
  • 1980
  • 1979
  • 1978
  • 1977
  • 1976
  • 1975
  • 1974
  • 1973
  • 1972
  • 1971
  • 1970
  • 1969
  • 1968
  • 1967
  • 1966
  • 1965
  • 1964
  • 1963
  • 1958
  • More
Subject
News
August 25, 2026
Scientists Develop Algorithm for More Reliable Processors in Data Centres
Researchers from HSE MIEM and Samara University have developed the LRF-3D algorithm to automatically bypass idle nodes in three-dimensional networks-on-chip. Thanks to its hierarchical architecture, the algorithm outperforms existing solutions in both speed and path accuracy, improving processor reliability for use in data centres, supercomputers, and AI computing. The source code and test results are publicly available.
August 24, 2026
Researchers Develop Method for Direct Generation of Regulatory DNA
Researchers at HSE University have developed a model for generating promoters and enhancers—DNA sequences that regulate gene activity. The model works directly with DNA nucleotides, without first transforming them into a continuous numerical representation. This solution could be useful for applications in synthetic biology and gene therapy. The study results were presented at the ICLR 2026 Workshop ‘Generative AI in Genomics (Gen^2): Barriers and Frontiers.’
August 21, 2026
Social Integration: At the Crossroads of Knowledge and Values
The International Laboratory for Social Integration Research (ILSIR) at HSE University studies the challenges faced by vulnerable groups and explores ways to help them participate fully in everyday life. To develop effective solutions, the laboratory’s researchers combine cutting-edge methods with practical fieldwork. In this interview with the HSE News Service, Laboratory Head Elena Iarskaia-Smirnova discusses the laboratory’s work.

 

Have you spotted a typo?
Highlight it, click Ctrl+Enter and send us a message. Thank you for your help!

Publications
  • Books
  • Articles
  • Chapters of books
  • Working papers
  • Report a publication
  • Research at HSE

?

Parse thicket representations of text paragraphs

P. 239–255.
Galitsky B., Ilvovsky D., Kuznetsov S., Strok F. V.

We develop a graph representation and learning technique for parse structures
for sentences and paragraphs of text. We introduce parse thicket
as a set of syntactic parse trees augmented by a number of arcs for intersentence
word-word relations such as coreference and taxonomies. These
arcs are also derived from other sources, including Rhetoric Structure and
Speech Act theory. We introduce respective indexing rules that identify inter-
sentence relations and join phrases connected by these relations in the
search index. We propose an algorithm for computing parse thickets from
parse trees. We develop a framework for automatic building and generalizing
of parse thickets. The proposed approach is used for evaluation in the
product search where search queries include multiple sentences. We draw
the comparison for search relevance improvement by pair-wise sentence
generalization and thicket-level generalization.

Language: English
Full text
Text on another site
Keywords: syntactic generalizationsearch relevanceparse thicketquestion answering
Publication based on the results of:
Mathematical Models, Algorithms, and Software Tools for the Intelligent Analysis of Big Textual and Structural Data (2013)

In book

Компьютерная лингвистика и интеллектуальные технологии: По материалам ежегодной Международной конференции «Диалог» (Бекасово, 29 мая - 2 июня 2013 г.). В 2-х т.
Т. 1: Основная программа конференции. Вып. 12 (19). , М.: РГГУ, 2013.
Similar publications
RuCLEVR: A Russian Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning
Biryukova K., Chelnokova D., Erkenova J. et al., Communications in Computer and Information Science 2024 Vol. 2364 CCIS P. 109 – 121
Added: February 25, 2026
CausalQA: A Benchmark for Causal Question Answering
Bondarenko A., Wolska M., Heindorf S. et al., , in: Proceedings of the 29th International Conference on Computational Linguistics.: International Committee on Computational Linguistics, 2023. P. 3296–3308.
Added: August 14, 2023
Relying on Discourse Trees to Extract Medical Ontologies from Text
Galitsky B., Ilvovsky D., Goncharova E., , in: Artificial Intelligence. RCAI 2021. Lecture Notes in Computer ScienceVol. 12948.: Springer, 2021. P. 215–231.
Added: October 28, 2021
DaNetQA: a yes/no Question Answering Dataset for the Russian Language
Glushkova T., Machnev A., Fenogenova A. et al., , in: Analysis of Images, Social Networks and Texts: 9th International Conference, AIST 2020, Skolkovo, Moscow, Russia, October 15–16, 2020, Revised Selected PapersVol. 12602.: Springer, 2021. P. 57–68.
Added: November 22, 2020
Обогащение контекста вопросов знаниями из ConceptNet для улучшения точности ответов
Smirnov D., Ilvovsky D., В кн.: Компьютерная лингвистика и интеллектуальные технологии: По материалам ежегодной международной конференции «Диалог» (Москва, 17 июня — 20 июня 2020 г.). Доклады студенческой сессии.: [б.и.], 2020..
Modern question answering models can achieve near-human accuracy of answers for factual questions about a given piece of text in English. In the meantime, such models fail to achieve the same performance on datasets of question, which require some background information, not presented in the question context. This paper describes experimental evaluation of simple question ...
Added: September 16, 2020
Russian Q&A Method Study: From Naive Bayes to Convolutional Neural Networks
Nikolaev K., Malafeev A., , in: Analysis of Images, Social Networks and Texts. 7th International Conference AIST 2018.: Springer, 2018. Ch. 12 P. 121–126.
This paper deals with automatic classification of questions in the Russian language. In contrast to previously used methods, we introduce a convolutional neural network for question classification. We took advantage of an existing corpus of 2008 questions, manually annotated in accordance with a pragmatic 14-class typology. We modified the data by reducing the typology to ...
Added: February 15, 2019
Russian-Language Question Classification: a New Typology and First Results
Nikolaev K., Malafeev A., , in: Analysis of Images, Social Networks and Texts. 6th International Conference, 2017, Revised Selected PapersVol. 10716.: Cham: Springer, 2018. Ch. 7 P. 72–81.
This paper deals with automatic classification of questions in the Russian language, a natural early step in building a question answering system. We developed a typology of Russian questions using interrogative particles, pronouns and word order as the main features. A corpus of 2008 questions was manually compiled and annotated according to our typology. We ...
Added: December 1, 2017
Выявление искаженной информации: подход с использованием дискурсивных связей
Galitsky B., Ilvovsky D., , in: Пятнадцатая национальная конференция по искусственному интеллекту с международным участием КИИ-2016 (3-7 октября 2016г., г.Смоленск, Россия): Труды конференцииТ. 1.: Смоленск: Универсум, 2016. P. 23–32.
A linguistic method for determining whether given text is a rumor or disinformation is proposed, based on web mining and linguistic technology comparing two text fragments. We hypothesize about a family of content generation algorithms which are capable of producing deception from a portion of genuine, original text. We then propose a disinformation detection algorithm ...
Added: March 5, 2017
Text integrity assessment: Sentiment profile vs rhetoric structure
Galitsky B., Ilvovsky D., Kuznetsov S., , in: Computational Linguistics and Intelligent Text Processing. 16th International Conference, CICLing 2015, Cairo, Egypt, April 14-20, 2015, Proceedings, Part II.Vol. 9042.: Berlin: Springer, 2015. P. 126–139.
We formulate the problem of text integrity assessment as learning thediscourse structure of text given the dataset of texts with high integrity and lowintegrity. We use two approaches to formalizing the discourse structures, sentimentprofile and rhetoric structures, relying on sentence-level sentiment classifierand rhetoric structure parsers respectively. To learn discourse structures, weuse the graph-based nearest neighbor ...
Added: November 7, 2015
Improving Text Retrieval Efficiency with Pattern Structures on Parse Thickets
Kuznetsov S., Strok F. V., Ilvovsky D. et al., , in: Proceedings of the Workshop Formal Concept Analysis Meets Information RetrievalVol. 977.: M.: CEUR Workshop Proceedings, 2013. P. 6–21.
We develop a graph representation and learning technique for parse  structures for paragraphs of text. We introduce Parse Thicket (PT) as a sum of  syntactic parse trees augmented by a number of arcs for inter-sentence word-word relations such as co-reference and taxonomic relations. These arcs are also derived from other sources, including Speech Act and ...
Added: November 18, 2013
Parse Thicket Representation for Multi-sentence Search
Galitsky B., Kuznetsov S., Usikov D., , in: Conceptual Structures for STEM Research and Education, 20th International Conference on Conceptual StructuresVol. 7735: Conceptual Structures for STEM Research and Education, 20th International Conference on Conceptual Structures.: Berlin, Heidelberg: Springer, 2013. P. 153–172.
We develop a graph representation and learning technique for parse structures for sentences and paragraphs of text. This technique is used to improve relevance answering complex questions where an answer is included in multiple sentences. We introduce Parse Thicket as a sum of syntactic parse trees augmented by a number of arcs for inter-sentence word-word ...
Added: June 2, 2013
  • About
  • About
  • Key Figures & Facts
  • Sustainability at HSE University
  • Faculties & Departments
  • International Partnerships
  • Faculty & Staff
  • HSE Buildings
  • HSE University for Persons with Disabilities
  • Public Enquiries
  • Studies
  • Admissions
  • Programme Catalogue
  • Undergraduate
  • Graduate
  • Exchange Programmes
  • Summer University
  • Summer Schools
  • Semester in Moscow
  • Business Internship
  • Research
  • International Laboratories
  • Research Centres
  • Research Projects
  • Monitoring Studies
  • Conferences & Seminars
  • Academic Jobs
  • Yasin (April) International Academic Conference on Economic and Social Development
  • Media & Resources
  • Publications by staff
  • HSE Journals
  • Publishing House
  • iq.hse.ru: commentary by HSE experts
  • Library
  • Economic & Social Data Archive
  • Video
  • HSE Repository of Socio-Economic Information
  • HSE1993–2026
  • Contacts
  • Copyright
  • Privacy Policy
  • Site Map
Edit