• A
  • A
  • A
  • АБВ
  • АБВ
  • АБВ
  • A
  • A
  • A
  • A
  • A
Обычная версия сайта
  • RU
  • EN
  • HSE University
  • Publications
  • Book chapter
  • Accelerating join of distributed datasets by a given criterion
  • RU
  • EN
Расширенный поиск
Высшая школа экономики
Национальный исследовательский университет
Priority areas
  • business informatics
  • economics
  • engineering science
  • humanitarian
  • IT and mathematics
  • law
  • management
  • mathematics
  • sociology
  • state and public administration
by year
  • 2028
  • 2027
  • 2026
  • 2025
  • 2024
  • 2023
  • 2022
  • 2021
  • 2020
  • 2019
  • 2018
  • 2017
  • 2016
  • 2015
  • 2014
  • 2013
  • 2012
  • 2011
  • 2010
  • 2009
  • 2008
  • 2007
  • 2006
  • 2005
  • 2004
  • 2003
  • 2002
  • 2001
  • 2000
  • 1999
  • 1998
  • 1997
  • 1996
  • 1995
  • 1994
  • 1993
  • 1992
  • 1991
  • 1990
  • 1989
  • 1988
  • 1987
  • 1986
  • 1985
  • 1984
  • 1983
  • 1982
  • 1981
  • 1980
  • 1979
  • 1978
  • 1977
  • 1976
  • 1975
  • 1974
  • 1973
  • 1972
  • 1971
  • 1970
  • 1969
  • 1968
  • 1967
  • 1966
  • 1965
  • 1964
  • 1963
  • 1958
  • More
Subject
News
September 11, 2026
How to Assess Students Knowledge in the Age of AI
A researcher at HSE University has proposed a flowchart to help lecturers decide how to assess students who use artificial intelligence. It shows where the use of AI should be restricted and where it can be incorporated into the learning process. The article has been published in IT Professional.
September 9, 2026
‘Balkan Hospitality Opens Doors: Studying Dialects on the Verge of Extinction
You cannot study spoken dialects from books. Instead, you need to go to a village, seek out its elders, and earn the trust of local residents before you can record hours of spontaneous stories. This is how Natalia Muravleva, Associate Professor at the Faculty of Humanities, conducts her research. Her internship in Serbia continued her long-standing study of dialects spoken by Macedonian settlers. In this interview, she discusses how diaspora cultural centres help researchers reach informants, why native speakers need to be interviewed only in their own language (otherwise, as she puts it, they may 'break'), and how a single field season helped her finalise her monograph. She also shares warm memories of autumn in Belgrade and of colleagues with whom grammar can be discussed in three languages at once.
September 9, 2026
Scientists Train Neural Network to Generate Process Plans from 3D Models
Researchers at the HSE FCS AI and Digital Science Institute have developed CAD2TechSpec, a framework that converts 3D models of mechanical parts into machining process plans—step-by-step instructions for machine tools. The solution aims to reduce the time required for the design and preparation of technical process documentation in mechanical engineering, aircraft manufacturing, and other high-tech industries. The study findings have been published in PeerJ Computer Science.

 

Have you spotted a typo?
Highlight it, click Ctrl+Enter and send us a message. Thank you for your help!

Publications
  • Books
  • Articles
  • Chapters of books
  • Working papers
  • Report a publication
  • Research at HSE

?

Accelerating join of distributed datasets by a given criterion

.
Tyryshkina Y.
Language: English
Keywords: MapReduceMapReduceanalytical systems

In book

Proceedings of 2022 IEEE Moscow Workshop on Electronic and Networking Technologies (MWENT)
Proceedings of 2022 IEEE Moscow Workshop on Electronic and Networking Technologies (MWENT)
M.: IEEE, 2022.
Similar publications
Triclustering in Big Data Setting
Egurnov D., Точилкин Д. С., Ignatov D. I., , in: Complex Data Analytics with Formal Concept Analysis.: Springer, 2022. P. 239–258.
In this paper, we describe versions of triclustering algorithms adapted for efficient calculations in distributed environments with MapReduce model or parallelisation mechanism provided by modern programming languages. OAC-family of triclustering algorithms shows good parallelisation capabilities due to the independent processing of triples of a triadic formal context. We provide time and space complexity of the ...
Added: November 1, 2022
Actor-oriented approach for business-process management of analytical system development
T.M. Pribylev, Zaytsev M. N., O.L. Vikentyeva, Proceedings of the Institute for System Programming of the RAS 2022 Vol. 34 No. 2 P. 57–65
This paper aims at investigating the feasibility of an actor-oriented approach for modelling analytical systems development business processes. The study analyzes existing management challenges of analytical systems development processes, identifies key business process modeling approaches, and proposes a modeling approach based on actor-oriented approach with high flexibility and enhanced control over business artifacts. The article also describes examples of ...
Added: September 5, 2022
Understanding join strategies in distributed systems
Tyryshkina Y., , in: International Seminar on Electron Devices Design and Production, SED 2021.: [б.и.], 2021.
In this paper, we consider the problem of reducing the cost of computer time by developing and implementing a method for accelerating the operation of connecting distributed data arrays according to a given criterion. The following tasks were solved: a study was conducted on the architecture of distributed data storages and parallel computing algorithms; on ...
Added: June 2, 2022
Метод ускорения объединения распределенных наборов данных по заданному критерию
Tyryshkina Y., Tumkovskiy S., Информационно-управляющие системы 2022 № 5(120) С. 2–11
Added: May 31, 2022
Ускорение объединения распределенных наборов данных по заданному критерию
Tyryshkina Y., В кн.: Международной научно-практической конференции «BIG DATA and Advanced Analytics».: Мн.: [б.и.], 2022.
В данной работе рассматривается проблема снижения затрат машинного времени за счет разработки и реализации метода ускорения операции соединения распределенных массивов данных по заданному критерию. Были решены следующие задачи: проведено исследование архитектуры распределенных хранилищ данных и алгоритмов параллельных вычислений; на основании этих исследований установлены лимитирующие стадии, замедляющие процесс переработки; разработан метод, исключающий установленные лимитирующие стадии; на ...
Added: May 31, 2022
Method for accelerating the operation of joining distributed datasets by a given criterion
Tyryshkina Y., , in: Международная научнопрактическая конференция «Информационные Инновационные Технологии», 2022.: [б.и.], 2022.
In this paper, we consider the problem of reducing the cost of computer time by developing and implementing a method for accelerating the operation of connecting distributed data arrays according to a given criterion. The following tasks were solved: a study was conducted on the architecture of distributed data storages and parallel computing algorithms; on ...
Added: May 31, 2022
Simplified Mapreduce Mechanism for Large Scale Data Processing
Ahmed Munna M. T., International Journal of Engineering and Technology 2018 Vol. 7 No. 8 P. 16–21
MapReduce has become a popular programming model for processing and running large-scale data sets with a parallel, distributed paradigm on a cluster. Hadoop MapReduce is needed especially for large scale data like big data processing. In this paper, we work to modify the Hadoop MapReduce Algorithm and implement it to reduce processing time. ...
Added: October 29, 2019
Распределенные горизонтально масштабируемые решения для управления данными
С.Д. Кузнецов, Посконин А. В., Труды Института системного программирования РАН 2013 Т. 24 С. 327–258
Many modern applications (such as large-scale Web-sites, social networks, research projects, business analytics, etc.) have to deal with very large data volumes (also referred to as “big data”) and high read/write loads. These applications require underlying data management systems to scale well in order to accommodate data growth and increasing workloads. High throughput, low latencies ...
Added: January 30, 2018
Большие данные: современные подходы к хранению и обработке
Клеменков П. А., Kuznetsov S. D., Труды Института системного программирования РАН 2012 Т. 23 С. 143–158
Big data challenged traditional storage and analysis systems in several new ways. In this paper we try to figure out how to overcome this challenges, why it's not possible to make it efficiently and describe three modern approaches to big data handling: NoSQL, MapReduce and real-time stream processing. The first section of the paper is ...
Added: October 31, 2017
Gomapreduce parallel computing model implementation on a cluster of plan9 virtual machines
Leokhin, Y., Myagkov, A., Panfilov, P., , in: 26th DAAAM International Symposium on Intelligent Manufacturing and Automation 2015Vol. 1.: NY: Curran Associates, Inc., 2015. P. 0656 – 0662.
In this paper, we present results of a computational evaluation of goMapReduce parallel programming model approach for solving distributed data processing problems. In some applications, particularly data center problems, including text processing the programming models can aggregate significant number of parallel processes. We first discuss the implementation of these approaches using both Linux and Plan9 ...
Added: November 26, 2016
Applying MapReduce to Conformance Checking
Shugurov I., Mitsyuk A. A., Proceedings of the Institute for System Programming of the RAS 2016 Vol. 28 No. 3 P. 103–122
Process mining is a relatively new research field, offering methods of business processes analysis and improvement, which are based on studying their execution history (event logs). Conformance checking is one of the main sub-fields of process mining. Conformance checking algorithms are aimed to assess how well a given process model, typically represented by a Petri ...
Added: September 12, 2016
Putting OAC-triclustering on MapReduce
Зудин С., Gnatyshak D. V., Ignatov D. I., , in: Proceedings of the Twelfth International Conference on Concept Lattices and Their Applications Clermont-Ferrand, France, October 13-16, 2015Vol. 1466.: Clermont-Ferrand: CEUR Workshop Proceedings, 2015. P. 47–58.
In our previous work an efficient one-pass online algorithm for triclustering of binary data (triadic formal contexts) was proposed. This algorithm is a modified version of the basic algorithm for OAC-triclustering approach; it has linear time and memory complexities. In this paper we parallelise it via map-reduce framework in order to make it suitable for big datasets. The results of ...
Added: October 23, 2015
Реализация модели параллельных вычислений GoMapReduce на операционной системе Plan9
Леохин Ю.Л., Мягков А.С., Информатизация образования и науки 2014 Т. 24 № 4 С. 111–118
In the article the implementation of the parallel programming model MapReduce, which is used in distributed computation, is considered. The results of the scalability research of the implementation running on the Plan9 operation system are shown. In addition, we have pointed out the main lines of the selected prototype development. ...
Added: October 23, 2014
  • About
  • About
  • Key Figures & Facts
  • Sustainability at HSE University
  • Faculties & Departments
  • International Partnerships
  • Faculty & Staff
  • HSE Buildings
  • HSE University for Persons with Disabilities
  • Public Enquiries
  • Studies
  • Admissions
  • Programme Catalogue
  • Undergraduate
  • Graduate
  • Exchange Programmes
  • Summer University
  • Summer Schools
  • Semester in Moscow
  • Business Internship
  • Research
  • International Laboratories
  • Research Centres
  • Research Projects
  • Monitoring Studies
  • Conferences & Seminars
  • Academic Jobs
  • Yasin (April) International Academic Conference on Economic and Social Development
  • Media & Resources
  • Publications by staff
  • HSE Journals
  • Publishing House
  • iq.hse.ru: commentary by HSE experts
  • Library
  • Economic & Social Data Archive
  • Video
  • HSE Repository of Socio-Economic Information
  • HSE1993–2026
  • Contacts
  • Copyright
  • Privacy Policy
  • Site Map
Edit