• A
  • A
  • A
  • АБВ
  • АБВ
  • АБВ
  • A
  • A
  • A
  • A
  • A
Обычная версия сайта
  • RU
  • EN
  • HSE University
  • Publications
  • Articles
  • When to Switch: Planning and Learning for Partially Observable Multi-Agent Pathfinding
  • RU
  • EN
Расширенный поиск
Высшая школа экономики
Национальный исследовательский университет
Priority areas
  • business informatics
  • economics
  • engineering science
  • humanitarian
  • IT and mathematics
  • law
  • management
  • mathematics
  • sociology
  • state and public administration
by year
  • 2028
  • 2027
  • 2026
  • 2025
  • 2024
  • 2023
  • 2022
  • 2021
  • 2020
  • 2019
  • 2018
  • 2017
  • 2016
  • 2015
  • 2014
  • 2013
  • 2012
  • 2011
  • 2010
  • 2009
  • 2008
  • 2007
  • 2006
  • 2005
  • 2004
  • 2003
  • 2002
  • 2001
  • 2000
  • 1999
  • 1998
  • 1997
  • 1996
  • 1995
  • 1994
  • 1993
  • 1992
  • 1991
  • 1990
  • 1989
  • 1988
  • 1987
  • 1986
  • 1985
  • 1984
  • 1983
  • 1982
  • 1981
  • 1980
  • 1979
  • 1978
  • 1977
  • 1976
  • 1975
  • 1974
  • 1973
  • 1972
  • 1971
  • 1970
  • 1969
  • 1968
  • 1967
  • 1966
  • 1965
  • 1964
  • 1963
  • 1958
  • More
Subject
News
October 1, 2026
HSE Researchers Show How Congenital Motor Disorders Affect Brain Development
Researchers from HSE University’s Institute for Cognitive Neuroscience have synthesised the findings of their previous studies on brain development in children with obstetric brachial plexus palsy and arthrogryposis. Their analysis shows that impaired motor function in early childhood not only limits children’s motor experience but also affects memory, categorical thinking, and information processing. The study has been published in Frontiers in Psychology.
October 1, 2026
Window into the Body: Scientists Develop Neural Network to Detect Risk of 15 Diseases from Retinal Images
Russian universities, with the participation of HSE University, Sber, and Z-union, have developed a neural network that can simultaneously assess the risk of 15 types of pathology from retinal photographs, including not only eye diseases but also cardiovascular conditions. The AI system can help clinicians detect potentially concerning changes at an early stage, identify signs reflecting the condition of retinal blood vessels, and determine whether a patient may need further examination. The paper has been published in Frontiers in Medicine.
September 30, 2026
'We Did Not Limit the Time for Questions'
The International Laboratory for Supercomputer Atomistic Modelling and Multi-Scale Analysis at HSE University held a major conference on molecular dynamics. Participants had the opportunity to attend all the presentations, while speakers were given as much time as they needed to answer questions. The HSE News Service interviewed Grigory Smirnov, Head of the Laboratory, and Genri Norman, Chief Research Fellow, about the conference preparations and the discussions it generated.

 

Have you spotted a typo?
Highlight it, click Ctrl+Enter and send us a message. Thank you for your help!

Publications
  • Books
  • Articles
  • Chapters of books
  • Working papers
  • Report a publication
  • Research at HSE

?

When to Switch: Planning and Learning for Partially Observable Multi-Agent Pathfinding

IEEE Transactions on Neural Networks and Learning Systems. 2024. Vol. 35. No. 12. P. 17411–17424.
Skrynnik A., Andreychuk A., Yakovlev K., Panov A.

Multi-agent pathfinding (MAPF) is a problem that involves finding a set of non-conflicting paths for a set of agents confined to a graph. In this work, we study a MAPF setting, where the environment is only partially observable for each agent, i.e., an agent observes the obstacles and other agents only within a limited field-of-view. Moreover, we assume that the agents do not communicate and do not share knowledge on their goals, intended actions, etc. The task is to construct a policy that maps the agent’s observations to actions. Our contribution is multifold. First, we propose two novel policies for solving partially observable MAPF (PO-MAPF): one based on heuristic search and another one based on reinforcement learning (RL). Next, we introduce a mixed policy that is based on switching between the two. We suggest three different switch scenarios: the heuristic, the deterministic, and the learnable one. A thorough empirical evaluation of all the proposed policies in a variety of setups shows that the mixing policy demonstrates the best performance is able to generalize well to the unseen maps and problem instances, and, additionally, outperforms the state-of-the-art counterparts (Primal2 and PICO). The source-code is available at https://github.com/AIRI-Institute/when-to-switch.

Research target: Computer Science
Language: English
DOI
Text on another site
Keywords: multiagent planningdeep reinforcement learningMulti-agent pathfinding
Similar publications
Нижние множества и свойства замкнутости классов функций подсчета
Ivanashev Y., Доклады Российской академии наук. Математика, информатика, процессы управления (ранее - Доклады Академии Наук. Математика) 2026 Т. 529 С. 93–101
Язык L является нижним для релятивизируемого сложностного класса C, если CL=C. Для классов #P, GapP и SpanP известны точные нижние классы языков: Low(#P) = UP ∩ coUP, Low(GapP) = SPP и Low(SpanP) = NP ∩ coNP. В этой статье мы доказываем, что Low(TotP) = P, и приводим характеризации нижних классов функций для #P, GapP, TotP ...
Added: September 28, 2026
Role of dislocations in the mobility of pinned helium bubbles: Molecular dynamics simulations in aluminum
Piliugin L., Antropov A., Lobashev E. et al., Journal of Nuclear Materials 2026 Vol. 632 Article 156876
The effects of dislocations on the mobility of gas nanobubbles pinned to them are considered as novel unex- plored mechanisms of accelerated fission gas release and investigated using classical molecular dynamics of helium bubbles in FCC aluminum. Non-equilibrium methods are developed to calculate the mobility of a pinned bubble both along and across the dislocation ...
Added: September 28, 2026
MPI+OpenMP implementation of resolution-of-the-identity Hartree-Fock method exploiting permutational symmetry of three-center electron repulsion integrals
Kashpurovich I., Oleynichenko A., Stegailov V., Supercomputing Frontiers and Innovations 2026 Vol. 13 No. 1 P. 52–73
We report a high-performance implementation of the resolution-of-the-identity Hartree–Fock method that fully exploits the permutational symmetry of three-center electron repulsion integrals (ERI). The present implementation adopts a hybrid MPI+OpenMP parallelization strategy. Two different algorithmic approaches (with and without the pre-transformation of ERIs) are compared. A custom data layout introduced previously is employed. Designed to efficiently ...
Added: September 28, 2026
A Three-Party W-State Quantum Secret Sharing Protocol with X-Gate Encoding and Forbidden-Outcome Detection
Teregulov T., Loubenets E. R., / Series Quantum Physics "arXiv". 2026. No. 2609.31472.
We develop a new three-party quantum secret-sharing (QSS) protocol based on a three-qubit W state. This protocol encodes the secret-sequence bits using X gates and employs randomly selected Hadamard operations and measurement bases to generate information and security-test rounds. We evaluate the efficiency of the proposed protocol and analyze its security against an internal adversary ...
Added: September 28, 2026
Bytedance и Open Source - открытые проекты от разработчика TikTok
Silakov D., Системный администратор 2026 С. 84–89
Social media users rarely think about what lies behind the beautiful facade of activity feeds, teeming with photos and video stories. However, the widespread popularity of such platforms generates a huge amount of all sorts of content that needs to be stored, processed quickly, and displayed, and in the era of AI, it also needs ...
Added: September 28, 2026
Shape-aware deep learning for models of production
Prokhorov A., Wei Z., Sang H. et al., Journal of Productivity Analysis 2026 Vol. 65 P. 1–16
The stochastic frontier model (SFM) is widely employed in the analysis of productivity and efficiency, yet strict parametric forms, such as the Cobb-Douglas and Translog functions, are often assumed for modeling production, leading to potential misspecification issues. While semi- and nonparametric SFMs offer greater flexibility, they face challenges in imposing monotonicity and concavity to maintain ...
Added: September 28, 2026
Inverse quickest path problem on networks under weighted l_\infty norm
Qian X., Guan X., Zhang B. et al., Journal of Global Optimization 2026
Inverse quickest path problem on networks ...
Added: September 27, 2026
Navigating Complexity: Statistical Methods, Data Analysis, and Machine Learning for Actionable Insights
Switzerland: Springer Cham, 2026.
This volume gathers selected, peer-reviewed contributions presented at the 19th Conference of the International Federation of Classification Societies (IFCS 2026), held on 14–16 July 2026 in Milan, Italy. Reflecting the volume’s motto, Navigating Complexity – Statistical Methods, Data Analysis, and Machine Learning for Actionable Insights, the papers showcase modern methodologies and real-world applications designed to extract ...
Added: September 25, 2026
An early warning system for emerging markets
Kraevskiy A., Sokolovskiy E., Prokhorov A., Emerging Markets Review 2026 No. 74 P. 1–19
Financial markets of emerging economies are vulnerable to extreme and cascading information spillovers, surges, sudden stops and reversals. With this in mind, we develop a new online early warning system (EWS) to detect what is referred to as ‘concept drift’ in machine learning, as a ‘regime shift’ in economics and as a ‘change-point’ in statistics. ...
Added: September 25, 2026
Экспериментальное сравнение HTTP/2 и HTTP/3 в условиях программно моделируемой сетевой деградации
Дубич Е. В., Schagin D., Славянский форум 2026 № 2 (52) С. 560–565
The paper compares HTTP/2 and HTTP/3 for static resource transfer under software-simulated network degradation. The experiment shows that HTTP/3 is not universally faster, but it is more stable as latency and packet loss increase. ...
Added: September 25, 2026
Synthesis of Acyclic Models for Processes Without Repeating Events
Joulitov A.K., Lomazova I.A., Proceedings of the Institute for System Programming of the RAS 2026 Vol. 38 No. 4(2) P. 215–224
In process mining, DFG (Directly-Follows Graph) models are popular due to their simplicity and clarity. However, if a process is acyclic but contains concurrent events, standard algorithms for discovering DFG models can generate "fake" cycles that do not actually exist in the event log. These cycles hinder the analysis of information processes, significantly reducing the ...
Added: September 24, 2026
Анализ протокола выработки общего ключа для управления микросхемой интеллектуальной карты
Добрина Д. Н., Nesterenko A., Прикладная дискретная математика. Приложение 2026 № 19 С. 151–159
Работа содержит результаты формального анализа криптографических механизмов, входящих в состав проекта методических рекомендаций «Защищенный универсальный протокол передачи данных и управления микросхемой интеллектуальной карты» (протокол SECUNDA). Получена формальная модель и перечень трудноразрешимых математических задач, трудоёмкостью решения которых можно оценить стойкость используемых криптографических механизмов. ...
Added: September 24, 2026
Discovering object-centric Petri nets with parametric arcs
I.I. Sergeev, I.A. Lomazova, Modeling and Analysis of Information Systems 2026 Vol. 33 No. 3 P. 394–419
Object-centric process mining has emerged as a powerful paradigm for analyzing event data involving multiple interacting business objects. Existing discovery techniques often rely on object-centric Petri nets with fixed arc multiplicities, limiting their ability to represent parametric resource consumption and production patterns and to capture quantitative dependencies between interacting object types. In this paper, we ...
Added: September 24, 2026
Hybrid Graph Retrieval-Augmented Language Agents for Collaborative Recommendation
Ivan Bulychev, Savchenko A., AI 2026 Vol. 7 No. 9 Article 380
Recent advances in large language model (LLM) agents have shown promise for autonomous decision-making in recommender systems. However, existing approaches suffer from two fundamental limitations: flat agent memories that conflate different information modalities and prohibitive computational costs that prevent scaling beyond a few hundred users. We propose Hybrid-GraphRAG, a recommender system that integrates hierarchical agent ...
Added: September 24, 2026
Risks and the image of the future in the study of AI technologies prospects
Snegirev A., Sychev S., Futures 2026 Vol. 183 P. 1–22
This study addresses the systemic identification and categorization of risks associated with AI development, arising from tensions between technological evolution and institutional, infrastructural, and economic contexts. Drawing on a constructionist methodology, we interpret technological risks as constitutive elements of expert communities' images of the future. Through in-depth interviews with 100 AI experts, proportionally representing corporate, ...
Added: September 23, 2026
Choosing Between AI Responses: How Valence, Arousal, and Dominance Shape User Preference
Parshakov P., Paklina S., International Journal of Human-Computer Interaction 2026 P. 1–17
This study examines how emotional tone shapes user preference in human–large language model (LLM) interaction. Drawing on the Computers as Social Actors framework, we treat conversational AI as a social communicator whose affective cues influence user judgments. Using large-scale pairwise preference data from LMSYS Chatbot Arena, we model emotional tone through the Valence–Arousal–Dominance framework and ...
Added: September 23, 2026
A Two-Stage Deep Reinforcement Learning Framework for Radio Resource Management and Network Slicing in 5G Heterogeneous Networks
Andrabi U., Wadood E., Ojha S. K. et al., IEEE Access 2026 Vol. 14 P. 103358–103375
The emergence of 5G networks, aimed at accommodating diverse service requirements such as enhanced Mobile Broadband (eMBB), Ultra-Reliable Low-Latency Communication (URLLC), and massive Machine-Type Communication (mMTC), has presented significant challenges in radio resource management and network slicing. In dynamic heterogeneous network systems, traditional heuristics and mathematical programming methods find it challenging to attain scalable multi-objective ...
Added: September 23, 2026
LLM-assisted writing and citation advantage: evidence from scientific publications before and after ChatGPT release
Paklina S., Parshakov P., Elena Rapoport, Scientometrics 2026 P. 1–26
Generative artificial intelligence has become a routine part of academic writing. While much of the debate has focused on questions of integrity and authorship, less attention has been paid to how AI-assisted writing may affect research evaluation itself. This paper asks a straightforward but important question: does the use of LLMs in academic writing change ...
Added: September 23, 2026
Label-Free Quantification in the Crux Toolkit
Kertesz-Farkas A., Acquaye F. L., Journal of Proteome Research 2026 Vol. 25 P. 3764–3768
Ultimately, most tandem mass spectrometry (MS/MS) proteomics experiments aim to not just detect but also quantify the proteins in a given complex sample. Here, we describe an extension to the Crux MS/MS analysis toolkit to enable label-free quantification of peptides. We demonstrate that Crux’s new quantification command, which is modeled after the algorithms implemented in ...
Added: September 23, 2026
Теория и практика совершенствования отечественных систем связи, автоматизации и информационной безопасности
СПб.: ВКАС им. Буденного, 2026.
В сборнике представлены статьи о результатах научных исследований и инновационных технических, технологических и методологических решений в области совершенствованию отечественных систем связи, автоматизации и информационной безопасности. ...
Added: September 23, 2026
Learning-Based UAV–RIS Secure Communication Under Eavesdropper Location Uncertainty
Ehab S. Suleiman, Ali J. Dayoub, , in: Proceedings of the 2026 8th International Youth Conference on Radio Electronics, Electrical and Power Engineering (REEPE).: IEEE, 2026. Ch. 165 P. 1–6.
Unmanned aerial vehicle (UAV)-assisted reconfigurable intelligent surface (RIS) systems can enhance physical layer security through joint mobility and propagation control. However, most existing designs assume the availability of the eavesdropper's channel state information (CSI), which is unrealistic in passive eavesdropping scenarios. In this paper, secure UAV-RIS downlink communication is studied under bounded eavesdropper location uncertainty, ...
Added: April 30, 2026
Навигация группы взаимозаменяемых агентов в непрерывной среде
Микрюкова А. В., Dergachev S., В кн.: XXII национальная конференция по искусственному интеллекту с международным участием (КИИ-2025)Т. 2.: СПб.: Санкт-Петербургский Федеральный исследовательский центр РАН, 2025. С. 183–194.
В данной работе рассматривается задача планирования путей для группы взаимозаменяемых агентов в непрерывной среде. В отличие от классического случая, в рассматриваемой постановке целевые позиции не закреплены за конкретными агентами. Проведён обзор современных методов много-агентного планирования, в ходе которого выделен алгоритм GAP, который использует непрерывное представление пространства и предъявляет минимальные требования к входным данным. Алгоритм был ...
Added: March 2, 2026
Optical stabilization for laser communication satellite systems through proportional–integral–derivative (PID) control and reinforcement learning approach
Бахшалиев Р. М., Reutov A., Vorobey S. et al., Review of Scientific Instruments 2025 Vol. 96 No. 3
One of the main issues of the satellite-to-ground optical communication, including free-space satellite quantum key distribution (QKD), is an achievement of the reasonable accuracy of positioning, navigation, and optical stabilization. Proportional–integral–derivative (PID) controllers can handle various control tasks in optical systems. Recent research shows the promising results in the area of composite control systems including ...
Added: May 13, 2025
Optimization of the Accelerator Control by Reinforcement Learning: A Simulation-Based Approach
Ibrahim A., Derkach D., Petrenko A. et al., Physics of Particles and Nuclei 2025 Vol. 56 No. 6 P. 1476–1481
Optimizing accelerator control is a critical challenge in experimental particle physics, requiring significant manual effort and resource expenditure. Traditional tuning methods are often time-consuming and reliant on expert input, highlighting the need for more efficient approaches. This study aims to create a simulation-based framework integrated with Reinforcement Learning (RL) to address these challenges. Using \texttt{Elegant} ...
Added: March 16, 2025
  • About
  • About
  • Key Figures & Facts
  • Sustainability at HSE University
  • Faculties & Departments
  • International Partnerships
  • Faculty & Staff
  • HSE Buildings
  • HSE University for Persons with Disabilities
  • Public Enquiries
  • Studies
  • Admissions
  • Programme Catalogue
  • Undergraduate
  • Graduate
  • Exchange Programmes
  • Summer University
  • Summer Schools
  • Semester in Moscow
  • Business Internship
  • Research
  • International Laboratories
  • Research Centres
  • Research Projects
  • Monitoring Studies
  • Conferences & Seminars
  • Academic Jobs
  • Yasin (April) International Academic Conference on Economic and Social Development
  • Media & Resources
  • Publications by staff
  • HSE Journals
  • Publishing House
  • iq.hse.ru: commentary by HSE experts
  • Library
  • Economic & Social Data Archive
  • Video
  • HSE Repository of Socio-Economic Information
  • HSE1993–2026
  • Contacts
  • Copyright
  • Privacy Policy
  • Site Map
Edit