• A
  • A
  • A
  • АБВ
  • АБВ
  • АБВ
  • A
  • A
  • A
  • A
  • A
Обычная версия сайта
  • RU
  • EN
  • HSE University
  • Publications
  • Articles
  • TreeDQN: Sample-efficient off-policy reinforcement learning for combinatorial optimization
  • RU
  • EN
Расширенный поиск
Высшая школа экономики
Национальный исследовательский университет
Priority areas
  • business informatics
  • economics
  • engineering science
  • humanitarian
  • IT and mathematics
  • law
  • management
  • mathematics
  • sociology
  • state and public administration
by year
  • 2028
  • 2027
  • 2026
  • 2025
  • 2024
  • 2023
  • 2022
  • 2021
  • 2020
  • 2019
  • 2018
  • 2017
  • 2016
  • 2015
  • 2014
  • 2013
  • 2012
  • 2011
  • 2010
  • 2009
  • 2008
  • 2007
  • 2006
  • 2005
  • 2004
  • 2003
  • 2002
  • 2001
  • 2000
  • 1999
  • 1998
  • 1997
  • 1996
  • 1995
  • 1994
  • 1993
  • 1992
  • 1991
  • 1990
  • 1989
  • 1988
  • 1987
  • 1986
  • 1985
  • 1984
  • 1983
  • 1982
  • 1981
  • 1980
  • 1979
  • 1978
  • 1977
  • 1976
  • 1975
  • 1974
  • 1973
  • 1972
  • 1971
  • 1970
  • 1969
  • 1968
  • 1967
  • 1966
  • 1965
  • 1964
  • 1963
  • 1958
  • More
Subject
News
September 11, 2026
How to Assess Students Knowledge in the Age of AI
A researcher at HSE University has proposed a flowchart to help lecturers decide how to assess students who use artificial intelligence. It shows where the use of AI should be restricted and where it can be incorporated into the learning process. The article has been published in IT Professional.
September 9, 2026
‘Balkan Hospitality Opens Doors: Studying Dialects on the Verge of Extinction
You cannot study spoken dialects from books. Instead, you need to go to a village, seek out its elders, and earn the trust of local residents before you can record hours of spontaneous stories. This is how Natalia Muravleva, Associate Professor at the Faculty of Humanities, conducts her research. Her internship in Serbia continued her long-standing study of dialects spoken by Macedonian settlers. In this interview, she discusses how diaspora cultural centres help researchers reach informants, why native speakers need to be interviewed only in their own language (otherwise, as she puts it, they may 'break'), and how a single field season helped her finalise her monograph. She also shares warm memories of autumn in Belgrade and of colleagues with whom grammar can be discussed in three languages at once.
September 9, 2026
Scientists Train Neural Network to Generate Process Plans from 3D Models
Researchers at the HSE FCS AI and Digital Science Institute have developed CAD2TechSpec, a framework that converts 3D models of mechanical parts into machining process plans—step-by-step instructions for machine tools. The solution aims to reduce the time required for the design and preparation of technical process documentation in mechanical engineering, aircraft manufacturing, and other high-tech industries. The study findings have been published in PeerJ Computer Science.

 

Have you spotted a typo?
Highlight it, click Ctrl+Enter and send us a message. Thank you for your help!

Publications
  • Books
  • Articles
  • Chapters of books
  • Working papers
  • Report a publication
  • Research at HSE

?

TreeDQN: Sample-efficient off-policy reinforcement learning for combinatorial optimization

Knowledge-Based Systems. 2026. Vol. 348. Article 116258.
Sorokin D., Kostin A., L. Savchenko, Gusev G., A.V. Savchenko

A convenient approach to optimally solving combinatorial optimization tasks is the Branch-and-Bound method.
Its branching heuristic can be learned to solve a large set of similar tasks. The promising results here are
achieved by the recently appeared on-policy reinforcement learning method based on the tree Markov Decision
Process. To overcome its main disadvantages, namely, very large training time and unstable training, we
propose TreeDQN (Tree Deep Q-Network), a sample-efficient off-policy RL method trained by optimizing the
geometric mean of expected return. To theoretically support the training procedure for our method, we prove
the contraction property of the Bellman operator for the tree MDP. As a result, our method requires up to
10 times less training data and performs faster than known on-policy methods on synthetic tasks. Moreover,
TreeDQN significantly outperforms the state-of-the-art techniques on a challenging practical task from the
ML4CO competition.

Research target: Computer Science
Language: English
Full text
DOI
Text on another site
Keywords: Марковский процессобучение с подкреплениемMixed integer linear programsReinforcement learningTree Markov Decision ProcessBellman operator’s contractionсмешанные целочисленные линейные программы
Similar publications
Метод кодирования голосового источника турбулентного типа на основе гибридной модели линейного предсказания
Савченко В. В., Savchenko L., Измерительная техника 2026 Т. 75 № 3 С. 105–113
Within the framework of a current area of research in the field of speech acoustics – non-invasive analysis of speech production processes – the acute problem of insufficient accuracy of parametric methods for coding a turbulent (noise) type voice source is considered. In order to overcome this problem, a method for coding a sound source with increased ...
Added: September 11, 2026
What Do Text-to-Image Models Know About the Languages of the World?
Фирсанова В. И., Journal of Mathematical Sciences 2024 Vol. 285 No. 1 P. 112–125
Text-to-image models use user-generated prompts to produce images. Such text-to-image models as DALL-E 2, Imagen, Stable Diffusion, and Midjourney can generate photorealistic or similar to human-drawn images. Apart from imitating human art, large text-to-image models have learned to produce combinations of pixels reminiscent of captions in natural languages. For example, a generated image might contain ...
Added: September 9, 2026
Разработка интерфейса виртуального ассистента преподавателя на основе технологий вызова функций и инженерии инструкций для больших языковых моделей
Фирсанова В. И., Человек: образ и сущность. Гуманитарные аспекты 2025 Vol. 2 No. 62 P. 203–214
Abstract. The paper highlights prompt engineering in academic setting to reduce plagiarism and increase students' interest. The research problem is the lack of a unified methodology for using artificial intelligence in education. The paper aims to create a generative artificial intelligence user interface, the Virtual Teaching Assistant. Teachers were interviewed, the first collection of presets ...
Added: September 9, 2026
Анализ согласованности голосования стран ЕАЭС и ОДКБ в ГА ООН с помощью иерархической кластеризации
Вохминцев И. В., Вестник международных организаций: образование, наука, новая экономика 2026 Т. 21 № 2
The EAEU and the CSTO are Russia’s principal regional international organisations. Understanding, assessing, and analysing the foreign-policy positions of the countries that belong to them is a matter of the state’s national interests. This determines the purpose of the study: to identify the level and the form of cohesion in the voting of EAEU and ...
Added: September 7, 2026
Pupillometry and autonomic nervous system responses to cognitive load and false feedback: an unsupervised machine learning approach
Alshanskaia E., Portnova G., Liaukovich K. et al., Frontiers in Neuroscience 2024 Vol. 18
Added: September 7, 2026
Oil Spill Segmentation in SAR Data Using ViT-UNet: Performance and Practical Insights
Зуенко Д. О., Trofimova E., Хайдарова И., IEEE Access 2026 Vol. 14 P. 121339–121357
Oil spill segmentation in Synthetic Aperture Radar (SAR) images is limited by noisy annotations in publicly available datasets and by architectural choices that interact with label quality in opposing directions. First, we introduce a manually refined version of the Deep-SAR Oil Spill (SOS) dataset, in which 36.25% of masks are corrected for false positives, missed ...
Added: September 7, 2026
Variational representation of weighted divergencies and error exponent function
Kelbert M., Statistics 2026 Vol. 60
We present variational representations for the weighted divergencies and exponential error function, and discuss implications for the statistical inference and entropic optimal transport. ...
Added: September 7, 2026
Scalable machine learning approach to disordered s-wave superconductors
Неверов В. Д., Красавин А. В., Vagov A. et al., Physical Review B: Condensed Matter and Materials Physics 2026 Vol. 113 P. 1–6
We develop a neural network approach to solve the self-consistent Bogoliubov-de Gennes equations in strongly disordered s-wave superconductors. The method accurately reproduces inhomogeneous gap distributions and generalizes to system sizes far larger than those used in training. It reduces computational scaling from O(N6 ) to O(N2), enabling quantitative analysis of percolation phenomena and the superconductor-insulator ...
Added: September 5, 2026
On the rate of Gaussian approximation for online linear regression problems
Sheshukova M., Durmus A., Khusainov M. et al., Statistics 2026 P. 1–25
In this paper, we consider the problem of Gaussian approximation for the online linear regression task. We derive the corresponding rates for the setting of a constant stepsize and study the explicit dependence of the convergence rate on the problem dimension d and quantities related to the design matrix. When the number of iterations n is known in advance, ...
Added: September 4, 2026
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence (UAI), PMLR Volume 337, 17-21 August 2026, KIT, Amsterdam, the Netherlands
Proceedings of Machine Learning Research , 2026.
Added: September 4, 2026
A unified frequency-domain framework for tilted slice localization and ischemic stroke detection
Khodadoust J., Kulikova S., Khodadoust F., Biomedical Signal Processing and Control 2027 Vol. 129 P. 111284–111284
Acute ischemic stroke (AIS) analysis from two-dimensional (2D) clinical imaging is hindered by uncontrolled slice tilt and geometric inconsistencies that violate the assumptions of pose-agnostic deep learning (DL) models. This paper proposes a unified geometry-aware, frequency-domain framework for tilted slice localization and ischemic stroke segmentation that explicitly decouples pose estimation from lesion analysis. The method ...
Added: September 2, 2026
Proceedings of the 2026 Fourth International Conference on Distributed Computing and High Performance Computing (DCHPC)
IEEE, 2026.
On behalf of the Organizing Committee, it is my great pleasure to extend a warm welcome to all participants of the Fourth International IEEE Conference on Distributed Computing and High-Performance Computing (DCHPC 2026), held in Tehran from May 10–11, 2026. This conference is jointly organized by the School of Computer Science at the Institute for Research in Fundamental Sciences (IPM) ...
Added: September 2, 2026
Discrete Markowitz Portfolio Optimization with Open-Source Classical and Quantum-Inspired Solvers: A Cross-Market Walk-Forward Study
S.M. Avdoshin, Patrushev K. A., Proceedings of the Institute for System Programming of the RAS 2026 Vol. 38 No. 4-2 P. 245–256
The cardinality-constrained Markowitz problem is NP-hard and traditionally solved with commercial MIQP solvers. Following the 2022 export restrictions that rendered both commercial MIQP software and cloud quantum platforms (IBM Quantum, D-Wave Leap) inaccessible from the Russian Federation, practitioners require open-source alternatives. This paper systematically compares three solver families for the discrete mean-variance problem: two open-source ...
Added: August 27, 2026
Benchmarking Synolitic Graphs for Autism Classification from Multisite Resting-State fMRI
Zaikin A., Vlasenko D., Zakharov D. et al., Diagnostics 2026 Vol. 16 No. 17 P. 1–15
Background/Objectives: Synolitic graphs (SGs) were developed for task-based fMRI, where edge weights encode the discriminative power of pairwise regional features; whether similar information can be recovered from resting-state data was untested. We benchmarked SGs for autism spectrum disorder (ASD) classification using the multisite ABIDE-I dataset (871 subjects: 403 subjects with ASD, 468 typical controls; 17 sites; CC200 atlas). Methods: Using ...
Added: August 27, 2026
Алгебра, теория чисел, дискретная геометрия и многомасштабное моделирование. Современные проблемы, приложения и проблемы истории. Материалы XXIV Международной конференции, посвящённой 110-летию со дня рождения академика Юрия Владимировича Линника и 110-летию со дня рождения профессора Андрея Борисовича Шидловского и 80-летию со дня рождения профессора Геннадия Ивановича Архипова
Тула: Тульский государственный педагогический университет им. Л.Н. Толстого, 2025.
Сборник содержит материалы, представленные на XXIV Международной конференции «Алгебра, теория чисел, дискретная геометрия и многомасштабное моделирование: современные проблемы, приложения и проблемы истории», посвящённой 110-летию со дня рождения академика Юрия Владимировича Линника и 110-летию со дня рождения профессора Андрея Борисовича Шидловского и 80-летию со дня рождения профессора Геннадия Ивановича Архипова. Материалы конференции будут полезны научным работникам, ...
Added: August 27, 2026
Characterizing the Scheduling Performance of 5G NR Base Stations Under Signaling and Data Traffic Constraints
Eduard Sopin, Nazarin A., Begishev V. et al., IEEE Transactions on Vehicular Technology 2026 Vol. 75 No. 6 P. 10995–11007
Aimed at rate-greedy applications having extreme requirements for the data rate at the air interface, 5G New Radio (NR) systems may experience problems when the number of user equipment (UE) in the coverage of the cell increases due to limited capacity of the physical downlink control channel (PDCCH).The aim of this study is to explore ...
Added: August 26, 2026
Improving Differential Equation Solving in Compact Language Models via Activation Steering and Reinforcement Learning
Surkov A., Ignatenko V., Koltsov S., Computers, Materials and Continua 2026 Vol. 88 No. 3 Article 74
Large language models have recently demonstrated promising capabilities in mathematical reasoning; however, their performance on tasks requiring strict symbolic manipulation, such as solving differential equations, remains limited, especially for compact models. In this work, we investigate whether activation steering combined with reinforcement learning can improve the quality of solutions generated by pretrained language models without ...
Added: July 8, 2026
UVIP: Model-Free Approach to Evaluate Reinforcement Learning Algorithms
Belomestny D., Levin I., Naumov A. et al., Journal of Optimization Theory and Applications 2026 Vol. 208 Article 89
Policy evaluation is an important instrument for the comparison of different algorithms in Reinforcement Learning (RL). However, even a precise knowledge of the value function Vπ corresponding to a policy π does not provide reliable information on how far the policy π is from the optimal one. We present a novel model-free upper value iteration ...
Added: February 10, 2026
Impact of self-learning based high-frequency traders on the stock market
Mansurov K., Semenov A., Dmitry Grigoriev et al., Expert Systems with Applications 2023 Vol. 232 Article 120567
In this paper we investigate the role of self-learning agents in multi-agent models of financial markets. We develop an agent-based simulation model of a financial market and, in addition to the agents with fixed strategies used in previous research, we introduce an agent with a self-learning strategy. To model the behavior of such an agent, ...
Added: July 11, 2025
Cryptocurrency Exchange Simulation
Mansurov K., Semenov A., Dmitry Grigoriev et al., Computational Economics 2024 Vol. 64 P. 2585–2603
In this paper, we consider the approach of applying state-of-the-art machine learning algorithms to simulate some financial markets. In this case, we choose the cryptocurrency market based on the assumption that such markets more active today. As a rule, they have more volatility, attracting riskier traders. Considering classic trading strategies, we also introduce an agent with a ...
Added: July 11, 2025
Optimal Approximation of Average Reward Markov Decision Processes
Sapronov Y., Yudin N., Computational Mathematics and Mathematical Physics 2025 Vol. 65 No. 3 P. 567–581
We continue to develop the concept of studying the ε-optimal policy for Average Reward Markov Decision Processes (AMDP) by reducing it to Discounted Markov Decision Processes (DMDP). Existing research often stipulates that the discount factor must not fall below a certain threshold. Typically, this threshold is close to one, and as is well-known, iterative methods ...
Added: June 10, 2025
The beer game bullwhip effect mitigation: a deep reinforcement learning approach
Rozhkov M., Alyamovskaya N., Zakhodiakin G., International Journal of Production Research 2025 Vol. 63 No. 18 P. 6630–6647
This article investigates the application of reinforcement learning (RL) methods to optimise a four-echelon linear supply chain model with stochastic demand. The proposed supply chain configuration is largely based on the production-distribution supply chain of the MIT Supply Chain Beer Game. We show that RL can significantly improve ordering efficiency and overall supply chain performance. ...
Added: March 24, 2025
Optimization of the Accelerator Control by Reinforcement Learning: A Simulation-Based Approach
Ibrahim A., Derkach D., Petrenko A. et al., Physics of Particles and Nuclei 2025 Vol. 56 No. 6 P. 1476–1481
Optimizing accelerator control is a critical challenge in experimental particle physics, requiring significant manual effort and resource expenditure. Traditional tuning methods are often time-consuming and reliant on expert input, highlighting the need for more efficient approaches. This study aims to create a simulation-based framework integrated with Reinforcement Learning (RL) to address these challenges. Using \texttt{Elegant} ...
Added: March 16, 2025
Компьютерное моделирование аффективных процессов в когнитивном контроле
Баланина С. Н., Berezner T., В кн.: Психология познания: материалы Всероссийской научной конференции. ЯрГУ, 6–8 декабря 2024 г. Материалы Всероссийской научной конференции памяти Дж. С. Брунера.: Яр.: ЯрГУ им. П. Г. Демидова, 2024. С. 45–48.
В настоящей работе мы предложили метод моделирования эмоциональной реакции, вызываемой стимулами в задаче Струпа. Наша модель отражает изменение валентности вызываемой реакции, то есть аффективной оценки стимула, по мере прохождения эксперимента. Мы использовали модель из класса алгоритмов обучения с подкреплением, разработанную Silvetti et al. (Silvetti et al., 2018). Результаты симуляции подтвердили, что вначале аффективная оценка выше ...
Added: December 28, 2024
  • About
  • About
  • Key Figures & Facts
  • Sustainability at HSE University
  • Faculties & Departments
  • International Partnerships
  • Faculty & Staff
  • HSE Buildings
  • HSE University for Persons with Disabilities
  • Public Enquiries
  • Studies
  • Admissions
  • Programme Catalogue
  • Undergraduate
  • Graduate
  • Exchange Programmes
  • Summer University
  • Summer Schools
  • Semester in Moscow
  • Business Internship
  • Research
  • International Laboratories
  • Research Centres
  • Research Projects
  • Monitoring Studies
  • Conferences & Seminars
  • Academic Jobs
  • Yasin (April) International Academic Conference on Economic and Social Development
  • Media & Resources
  • Publications by staff
  • HSE Journals
  • Publishing House
  • iq.hse.ru: commentary by HSE experts
  • Library
  • Economic & Social Data Archive
  • Video
  • HSE Repository of Socio-Economic Information
  • HSE1993–2026
  • Contacts
  • Copyright
  • Privacy Policy
  • Site Map
Edit