• A
  • A
  • A
  • АБВ
  • АБВ
  • АБВ
  • A
  • A
  • A
  • A
  • A
Обычная версия сайта
  • RU
  • EN
  • HSE University
  • Publications
  • Articles
  • Optimal Approximation of Average Reward Markov Decision Processes
  • RU
  • EN
Расширенный поиск
Высшая школа экономики
Национальный исследовательский университет
Priority areas
  • business informatics
  • economics
  • engineering science
  • humanitarian
  • IT and mathematics
  • law
  • management
  • mathematics
  • sociology
  • state and public administration
by year
  • 2028
  • 2027
  • 2026
  • 2025
  • 2024
  • 2023
  • 2022
  • 2021
  • 2020
  • 2019
  • 2018
  • 2017
  • 2016
  • 2015
  • 2014
  • 2013
  • 2012
  • 2011
  • 2010
  • 2009
  • 2008
  • 2007
  • 2006
  • 2005
  • 2004
  • 2003
  • 2002
  • 2001
  • 2000
  • 1999
  • 1998
  • 1997
  • 1996
  • 1995
  • 1994
  • 1993
  • 1992
  • 1991
  • 1990
  • 1989
  • 1988
  • 1987
  • 1986
  • 1985
  • 1984
  • 1983
  • 1982
  • 1981
  • 1980
  • 1979
  • 1978
  • 1977
  • 1976
  • 1975
  • 1974
  • 1973
  • 1972
  • 1971
  • 1970
  • 1969
  • 1968
  • 1967
  • 1966
  • 1965
  • 1964
  • 1963
  • 1958
  • More
Subject
News
August 25, 2026
Scientists Develop Algorithm for More Reliable Processors in Data Centres
Researchers from HSE MIEM and Samara University have developed the LRF-3D algorithm to automatically bypass idle nodes in three-dimensional networks-on-chip. Thanks to its hierarchical architecture, the algorithm outperforms existing solutions in both speed and path accuracy, improving processor reliability for use in data centres, supercomputers, and AI computing. The source code and test results are publicly available.
August 24, 2026
Researchers Develop Method for Direct Generation of Regulatory DNA
Researchers at HSE University have developed a model for generating promoters and enhancers—DNA sequences that regulate gene activity. The model works directly with DNA nucleotides, without first transforming them into a continuous numerical representation. This solution could be useful for applications in synthetic biology and gene therapy. The study results were presented at the ICLR 2026 Workshop ‘Generative AI in Genomics (Gen^2): Barriers and Frontiers.’
August 21, 2026
Social Integration: At the Crossroads of Knowledge and Values
The International Laboratory for Social Integration Research (ILSIR) at HSE University studies the challenges faced by vulnerable groups and explores ways to help them participate fully in everyday life. To develop effective solutions, the laboratory’s researchers combine cutting-edge methods with practical fieldwork. In this interview with the HSE News Service, Laboratory Head Elena Iarskaia-Smirnova discusses the laboratory’s work.

 

Have you spotted a typo?
Highlight it, click Ctrl+Enter and send us a message. Thank you for your help!

Publications
  • Books
  • Articles
  • Chapters of books
  • Working papers
  • Report a publication
  • Research at HSE

?

Optimal Approximation of Average Reward Markov Decision Processes

Computational Mathematics and Mathematical Physics. 2025. Vol. 65. No. 3. P. 567–581.
Sapronov Y., Yudin N.

We continue to develop the concept of studying the ε-optimal policy for Average Reward Markov Decision Processes (AMDP) by reducing it to Discounted Markov Decision Processes (DMDP). Existing research often stipulates that the discount factor must not fall below a certain threshold. Typically, this threshold is close to one, and as is well-known, iterative methods used to find the optimal policy for DMDP become less effective as the discount factor approaches this value.

Our work distinguishes itself from existing studies by allowing for inaccuracies in solving the empirical Bellman equation. Despite this, we have managed to maintain the sample complexity that aligns with the latest results. We have succeeded in separating the contributions from the inaccuracy of approximating the transition matrix and the residuals in solving the Bellman equation in the upper estimate so that our findings enable us to determine the total complexity of the epsilon-optimal policy analysis for DMDP across any method with a theoretical foundation in iterative complexity.

Research target: Mathematics Computer Science
Language: English
DOI
Text on another site
Keywords: Markov Decision Processesвычислительная сложностьобучение с подкреплениемадгритмы и алгоритмическая сложностьразмер выборкиsample complexityreinforcement learning (RL)iteration complexityмарковские процессы принятия решений
Similar publications
Discrete Markowitz Portfolio Optimization with Open-Source Classical and Quantum-Inspired Solvers: A Cross-Market Walk-Forward Study
Avdoshin S.M., Patrushev K. A., Proceedings of the Institute for System Programming of the RAS 2026 No. 4 часть 2 P. 245–256
The cardinality-constrained Markowitz problem is NP-hard and traditionally solved with commercial MIQP solvers. Following the 2022 export restrictions that rendered both commercial MIQP software and cloud quantum platforms (IBM Quantum, D-Wave Leap) inaccessible from the Russian Federation, practitioners require open-source alternatives. This paper systematically compares three solver families for the discrete mean-variance problem: two open-source ...
Added: August 27, 2026
Benchmarking Synolitic Graphs for Autism Classification from Multisite Resting-State fMRI
Zaikin A., Vlasenko D., Zakharov D. et al., Diagnostics 2026 Vol. 16 No. 17 P. 1–15
Background/Objectives: Synolitic graphs (SGs) were developed for task-based fMRI, where edge weights encode the discriminative power of pairwise regional features; whether similar information can be recovered from resting-state data was untested. We benchmarked SGs for autism spectrum disorder (ASD) classification using the multisite ABIDE-I dataset (871 subjects: 403 subjects with ASD, 468 typical controls; 17 sites; CC200 atlas). Methods: Using ...
Added: August 27, 2026
Pyramidal solitons: existence, instability, and interactions
Melnikov I., Pelinovsky E., Nonlinear Dynamics 2026 No. 114 Article 934
Pyramidal solitons (solitons with more than two inflection points) of the generalized Korteweg–de Vries (KdV) equation are investigated. Necessary and sufficient conditions for the existence of such structures are presented. Within the framework of the generalized Gardner equation — the simplest model that admits pyramidal solitons as solutions — it is shown that these solutions ...
Added: August 27, 2026
Алгебра, теория чисел, дискретная геометрия и многомасштабное моделирование. Современные проблемы, приложения и проблемы истории. Материалы XXIV Международной конференции, посвящённой 110-летию со дня рождения академика Юрия Владимировича Линника и 110-летию со дня рождения профессора Андрея Борисовича Шидловского и 80-летию со дня рождения профессора Геннадия Ивановича Архипова
Тула: Тульский государственный педагогический университет им. Л.Н. Толстого, 2025.
Сборник содержит материалы, представленные на XXIV Международной конференции «Алгебра, теория чисел, дискретная геометрия и многомасштабное моделирование: современные проблемы, приложения и проблемы истории», посвящённой 110-летию со дня рождения академика Юрия Владимировича Линника и 110-летию со дня рождения профессора Андрея Борисовича Шидловского и 80-летию со дня рождения профессора Геннадия Ивановича Архипова. Материалы конференции будут полезны научным работникам, ...
Added: August 27, 2026
О подходе к построению последовательности псевдослучайных чисел, основанном на разложениях 𝐸-функций с периодическими коэффициентами
Nesterenko A., Чирский В. Г., Матвеев В. Ю., Чебышевский сборник 2026 Т. 27 № 2 С. 118–127
The article presents the results of a practical study of the statistical properties of some sequences that are values ​​of functions of a special type. ...
Added: August 26, 2026
Characterizing the Scheduling Performance of 5G NR Base Stations Under Signaling and Data Traffic Constraints
Eduard Sopin, Nazarin A., Begishev V. et al., IEEE Transactions on Vehicular Technology 2026 Vol. 75 No. 6 P. 10995–11007
Aimed at rate-greedy applications having extreme requirements for the data rate at the air interface, 5G New Radio (NR) systems may experience problems when the number of user equipment (UE) in the coverage of the cell increases due to limited capacity of the physical downlink control channel (PDCCH).The aim of this study is to explore ...
Added: August 26, 2026
Генерация исходного кода с использованием больших языковых моделей: систематический обзор методологии Вайб-кодинг
Джонов А. Т., Avdoshin S. M., Информационные технологии 2026 Т. 32 № 8 С. 421–427
This systematic review presents an analysis of the "Vibe Coding" methodology — a contemporary approach to the iterative software development process using Large Language Models (LLMs). Code generation tools are transforming software development by enabling programmers to formulate tasks and describe the desired behavior of software in natural language, while LLMs generate source code corresponding ...
Added: August 25, 2026
Balanced sets and homotopy invariants of covers
Блудов М. В., Journal of Fixed Point Theory and Applications 2026 No. 28 Article 73
In this paper, we study a construction of homotopy invariants of open or closed covers, where the homotopy class is defined relative to a pair (V, r), with V a finite set of points in  and r a point in the interior of their convex hull. We show that the simplicial complex of non-balanced subsets associated with (V, r) has the homotopy ...
Added: August 25, 2026
An adaptive image watermarking scheme using cooperation of HBA and RSA metaheuristics
Melman A., Evsyutin O., Journal of the Franklin Institute 2026 Vol. 363 No. 15 Article 109005
Open access to images creates opportunities for violation of the authors' rights. Digital watermarks can be used to securely publish images online. They are invisibly added into the images before publication and can be extracted at any time to verify ownership. However, achieving a balance between embedding imperceptibility and robustness to image processing operations is ...
Added: August 25, 2026
Exact solutions for transport of distributed colloids in porous media
L.I. Kuzmina, Osipov , Y. V., Kolokoltseva, T. N., Advances in Water Resources 2026 Vol. 213 Article 105335
We study 1D transport of mobilized particles, detached from the solid matrix, in porous media (so-called fines migration). Three types of colloidal-suspension flow models are considered: (i) averaged model for multicomponent colloids with distributed properties; (ii) flow of binary colloids with interacting particles; (iii) discrete system for multicomponent low-concentration colloid. These models account for distributed ...
Added: August 25, 2026
MODEL OF DEEP BED FILTRATION IN A POROUS MEDIUM WITH HETEROGENEOUS POROSITY
Liudmila I. Kuzmina, Osipov, Y. V., International Journal for Computational Civil and Structural Engineering 2026 Vol. 22 No. 2 P. 138–148
Modeling the transport and sedimentation of small particles of suspensions and colloids in porous rocks is an important problem in subsurface hydromechanics. Particles entrained in fluid are transported and retained in the rock pores. The filtration process is determined by the number and size of pores and is characterized by porosity—the ratio of the void ...
Added: August 25, 2026
Exact Solutions to One-Dimensional Problem of Oil Displacement by Chemical
L.I.Kuzmina, Osipov Y. V., Mathematical notes, ISSN 0001-4346 2025 Vol. 116 No. 6 P. 1251–1261
We consider the displacement of oil by water with active chemical reagents in porous media. A one-dimensional model of reagent transport and deposition is defined by a hyperbolic system of first-order equations. The purpose of the article is to find the conditions for the existence of a continuous solution with discontinuous boundary and initial conditions. ...
Added: August 25, 2026
Proceedings of the 2026 12th International Conference on Control, Decision and Information Technologies (CoDIT) (Italy, Bari, July 13–16, 2026)
IEEE, 2026.
It is with great pleasure that we welcome all the participants of the 12th Conference on Control, Decision and Information Technologies (CoDIT 2026) at the Polytechnic University of Bari – Orabona Street 4, 70125 Bari, Italy, July 13-16, 2026. CoDIT has grown to become one of the largest conferences organized in Europe and in the ...
Added: August 24, 2026
From data to knowledge: artificial intelligence methods for studying comorbidity in electronic health records
Лукьяненко Д. В., Ragimova A., Мухорина А. et al., European Physical Journal: Special Topics 2026 P. 1–24
Electronic health records (EHRs) contain vast volumes of clinical information that encode complex relationships between diseases. Traditional approaches to the analysis of interrelated or co-occurring diseases have focused on pairwise associations between diagnoses, missing the higher-order structures that characterise multimorbid patients. The present paper offers a narrative review of existing statistical, machine-learning, and artificial intelligence ...
Added: August 20, 2026
Proceedings of the Generative Code Intelligence Workshop (GeCoIn 2026), co-located with the 35th International Joint Conference on Artificial Intelligence (IJCAI-ECAI 2026)
CEUR-WS.org, 2026.
The second edition of the Generative Code Intelligence Workshop (GeCoIn 2026) was held in conjunction with the 35th International Joint Conference on Artificial Intelligence (IJCAI-ECAI 2026), in Bremen, Germany, August 16, 2026. The workshop arose from the desire to bring together a research community that has, in recent years, witnessed rapid progress in the application ...
Added: August 20, 2026
MM-PSYCHE: Multimodal Multitask Psychological Characteristic Estimation Through Cross-Domain Semi-Supervised Learning
Ryumina E., Aksenov A., Koryakovskaya D. et al., IEEE Access 2026 Vol. 14 P. 124759–124778
Psychological characteristic estimation from multimodal in-the-wild behavior is usually studied using separate corpora, each annotated for a single target task. Such annotation fragmentation limits cross-task learning and cross-domain generalization across affective, dispositional, and interactional phenomena. To address this problem, we use emotion, apparent personality trait, and ambivalence recognition as representative tasks and introduce MM-PSYCHE, a ...
Added: August 20, 2026
Mathematical methods of reinforcement learning
Belomestny D., Gasnikov A., Gladin E. et al., Russian Mathematical Surveys 2026 Vol. 81 No. 4(490) P. 3–90
Reinforcement learning (RL) is increasingly grounded in tools from probability, optimization, and operator theory. This survey organizes the mathematical structures that underpin the design and analysis of modern algorithms in RL. We begin from Markov decision processes (MDPs) and the Bellman operators, emphasizing contraction mappings, monotonicity, and fixed-point theory that yield convergence guarantees and rates ...
Added: August 3, 2026
Optimal navigation in two-dimensional flows: Control theory and reinforcement learning
Parfenyev V., Physical Review E - Statistical, Nonlinear, and Soft Matter Physics 2026 Vol. 114 No. 1 Article 015104
Zermelo's navigation problem seeks the trajectory of minimal travel time between two points in a fluid flow. We address this problem for an agent -- such as a floating drone or active particle -- that is advected by a two-dimensional flow, self-propels at a fixed speed smaller than or comparable to the characteristic flow velocity, ...
Added: July 17, 2026
Improving Differential Equation Solving in Compact Language Models via Activation Steering and Reinforcement Learning
Surkov A., Ignatenko V., Koltsov S., Computers, Materials and Continua 2026 Vol. 88 No. 3 Article 74
Large language models have recently demonstrated promising capabilities in mathematical reasoning; however, their performance on tasks requiring strict symbolic manipulation, such as solving differential equations, remains limited, especially for compact models. In this work, we investigate whether activation steering combined with reinforcement learning can improve the quality of solutions generated by pretrained language models without ...
Added: July 8, 2026
TreeDQN: Sample-efficient off-policy reinforcement learning for combinatorial optimization
Sorokin D., Kostin A., Savchenko L. et al., Knowledge-Based Systems 2026 Vol. 348 Article 116258
A convenient approach to optimally solving combinatorial optimization tasks is the Branch-and-Bound method. Its branching heuristic can be learned to solve a large set of similar tasks. The promising results here are achieved by the recently appeared on-policy reinforcement learning method based on the tree Markov Decision Process. To overcome its main disadvantages, namely, very large training time ...
Added: June 10, 2026
Universal Comparison Methodology for Hough Transform Approaches
Kazimirov D., Vitalii Gulevskii, Kroshnin A. et al., Mathematics 2026 Article 1136
The Hough transform (HT) is widely used in computer vision, tomography, and neural networks. Numerous algorithms for HT computation have been proposed, making their systematic comparison essential. However, existing comparative methodologies are either non-universal and limited to certain HT formulations, or task-oriented, relying on application-specific criteria that do not fully capture algorithmic properties. This paper ...
Added: May 28, 2026
О СЛОЖНОСТИ ПРОБЛЕМЫ ТОТАЛЬНОЙ ВЫВОДИМОСТИ В НЕУКОРАЧИВАЮЩИХ И КОНТЕКСТНО-СВОБОДНЫХ ГРАММАТИКАХ
Dudakov S., Карлов Б. Н., Доклады Российской академии наук. Математика, информатика, процессы управления (ранее - Доклады Академии Наук. Математика) 2025 Т. 524 № 1 С. 11–18
In this paper we study the problem of total derivability in context-free, noncontracting, and context-sensitive grammars. Given a grammar and a terminal word, one has to determine whether there exists a derivation of this word which uses each production no less than a given number of times. It is proved that the problem of total ...
Added: March 18, 2026
О схлопывании вероятностных иерархий. I
Speranski S. O., Алгебра и логика 2013 Т. 52 № 2 С. 236–254
Изучаются иерархии проблем общезначимости для префиксных фрагментов вероятностной логики с кванторами по пропозициональным формулам, обозначаемой QPL, и её вариантов. Доказывается: если подполе F вещественных чисел определимо в стандартной модели арифметики посредством формулы второго порядка, не содержащей кванторов по множествам, то проблема общезначимости над F-значными вероятностными структурами для $\Sigma_4$-QPL-предложений является $\Pi^1_1$-полной и, как следствие, соответствующая иерархия проблем общезначимости схлопывается. Более того, при ...
Added: December 27, 2025
Некоторые классификации сложности задачи о вершинной 3-раскраске
Дахно Г. С., Malyshev D., Математические заметки 2026 Т. 119 № 3 С. 360–376
Наследственный класс — множество графов, замкнутое относительно удаления вершин. Каждый такой класс имеет каноническое описание посредством минимальных запрещенных порожденных фрагментов. Задача о вершинной 3-раскраске (задача 3-ВР) для заданного графа состоит в том, чтобы определить, а можно ли множество его вершин разбить на три подмножества попарно несмежных вершин. Известна дихотомия сложности этой задачи для всех наследственных ...
Added: November 26, 2025
  • About
  • About
  • Key Figures & Facts
  • Sustainability at HSE University
  • Faculties & Departments
  • International Partnerships
  • Faculty & Staff
  • HSE Buildings
  • HSE University for Persons with Disabilities
  • Public Enquiries
  • Studies
  • Admissions
  • Programme Catalogue
  • Undergraduate
  • Graduate
  • Exchange Programmes
  • Summer University
  • Summer Schools
  • Semester in Moscow
  • Business Internship
  • Research
  • International Laboratories
  • Research Centres
  • Research Projects
  • Monitoring Studies
  • Conferences & Seminars
  • Academic Jobs
  • Yasin (April) International Academic Conference on Economic and Social Development
  • Media & Resources
  • Publications by staff
  • HSE Journals
  • Publishing House
  • iq.hse.ru: commentary by HSE experts
  • Library
  • Economic & Social Data Archive
  • Video
  • HSE Repository of Socio-Economic Information
  • HSE1993–2026
  • Contacts
  • Copyright
  • Privacy Policy
  • Site Map
Edit