• A
  • A
  • A
  • АБВ
  • АБВ
  • АБВ
  • A
  • A
  • A
  • A
  • A
Обычная версия сайта
  • RU
  • EN
  • HSE University
  • Publications
  • Articles
  • Reinforcement Procedure for Randomized Machine Learning
  • RU
  • EN
Расширенный поиск
Высшая школа экономики
Национальный исследовательский университет
Priority areas
  • business informatics
  • economics
  • engineering science
  • humanitarian
  • IT and mathematics
  • law
  • management
  • mathematics
  • sociology
  • state and public administration
by year
  • 2028
  • 2027
  • 2026
  • 2025
  • 2024
  • 2023
  • 2022
  • 2021
  • 2020
  • 2019
  • 2018
  • 2017
  • 2016
  • 2015
  • 2014
  • 2013
  • 2012
  • 2011
  • 2010
  • 2009
  • 2008
  • 2007
  • 2006
  • 2005
  • 2004
  • 2003
  • 2002
  • 2001
  • 2000
  • 1999
  • 1998
  • 1997
  • 1996
  • 1995
  • 1994
  • 1993
  • 1992
  • 1991
  • 1990
  • 1989
  • 1988
  • 1987
  • 1986
  • 1985
  • 1984
  • 1983
  • 1982
  • 1981
  • 1980
  • 1979
  • 1978
  • 1977
  • 1976
  • 1975
  • 1974
  • 1973
  • 1972
  • 1971
  • 1970
  • 1969
  • 1968
  • 1967
  • 1966
  • 1965
  • 1964
  • 1963
  • 1958
  • More
Subject
News
September 25, 2026
AI Users Earn Up to 41.8% More Than Non-Users
Research conducted by economists at HSE University has revealed a significant correlation between the regular use of GenAI in the workplace and higher pay among Russian employees. The study found that individuals who frequently use GenAI in their professional activities earn notably more than those who reject these new tools or resort to them occasionally. The salary premium for highly qualified specialists reaches 41.8%. The article was published in the Voprosy Ekonomiki journal.
September 24, 2026
‘Feedback and Constructive Criticism Are Essential in Our Profession
Vincent Fardeau, Associate Professor at HSE ICEF, has reached a major career milestone: he recently published his paper ‘Asymmetric Thin Markets’ in the Journal of Financial Economics, successfully passed his major academic review, and received tenure. In this interview, Vincent discusses the story behind the paper, explains the concept of asymmetric thin markets, and shares his advice for young scholars aiming to publish in top-tier journals.
September 22, 2026
Personal Interest in Doctoral Thesis Topic Most Important for Confidence in Successful Defence
A researcher at HSE University analysed data on 1,539 doctoral students from 161 Russian universities to identify which features of a thesis topic are associated with academic success and engagement. The most important factor was found to be personal interest in the research topic, which was associated with almost all key aspects of doctoral programme experience—from engaging with the academic supervisor to research activity and confidence about successfully defending the thesis. The findings have been published in Higher Education.

 

Have you spotted a typo?
Highlight it, click Ctrl+Enter and send us a message. Thank you for your help!

Publications
  • Books
  • Articles
  • Chapters of books
  • Working papers
  • Report a publication
  • Research at HSE

?

Reinforcement Procedure for Randomized Machine Learning

Mathematics. 2023. Vol. 11. No. 17. Article 3651.
Yuri S. Popkov, Dubnov Y. A., Alexey Yu. Popkov

This paper is devoted to problem-oriented reinforcement methods for the numerical implementation of Randomized Machine Learning. We have developed a scheme of the reinforcement procedure based on the agent approach and Bellman’s optimality principle. This procedure ensures strictly monotonic properties of a sequence of local records in the iterative computational procedure of the learning process. The dependences of the dimensions of the neighborhood of the global minimum and the probability of its achievement on the parameters of the algorithm are determined. The convergence of the algorithm with the indicated probability to the neighborhood of the global minimum is proved.

Research target: Mathematics Computer Science
Language: English
Full text
DOI
Text on another site
Keywords: reinforcement learningBellman’s optimality principle randomized machine learning
Similar publications
Iterative construction of the R-matrices in arbitrary dimensions
Pyatov P. N., Pivovarov P. A., / Series math "arxiv.org". 2026. No. 2609.06274.
We investigate a special ansats that allows for an iterative solution of the constant Yang-Baxter equation. Testing this ansatz, we construct four sequences of the constant R-matrices. In each sequence the R-matrices act on the tensor squares of vector spaces of linearly growing dimensions. Each R-matrix also depends on a single complex parameter. By analyzing the ...
Added: September 24, 2026
Hybrid Graph Retrieval-Augmented Language Agents for Collaborative Recommendation
Ivan Bulychev, Savchenko A., AI 2026 Vol. 7 No. 9 Article 380
Recent advances in large language model (LLM) agents have shown promise for autonomous decision-making in recommender systems. However, existing approaches suffer from two fundamental limitations: flat agent memories that conflate different information modalities and prohibitive computational costs that prevent scaling beyond a few hundred users. We propose Hybrid-GraphRAG, a recommender system that integrates hierarchical agent ...
Added: September 24, 2026
Risks and the image of the future in the study of AI technologies prospects
Snegirev A., Sychev S., Futures 2026 Vol. 183 P. 1–22
This study addresses the systemic identification and categorization of risks associated with AI development, arising from tensions between technological evolution and institutional, infrastructural, and economic contexts. Drawing on a constructionist methodology, we interpret technological risks as constitutive elements of expert communities' images of the future. Through in-depth interviews with 100 AI experts, proportionally representing corporate, ...
Added: September 23, 2026
Choosing Between AI Responses: How Valence, Arousal, and Dominance Shape User Preference
Parshakov P., Paklina S., International Journal of Human-Computer Interaction 2026 P. 1–17
This study examines how emotional tone shapes user preference in human–large language model (LLM) interaction. Drawing on the Computers as Social Actors framework, we treat conversational AI as a social communicator whose affective cues influence user judgments. Using large-scale pairwise preference data from LMSYS Chatbot Arena, we model emotional tone through the Valence–Arousal–Dominance framework and ...
Added: September 23, 2026
A Two-Stage Deep Reinforcement Learning Framework for Radio Resource Management and Network Slicing in 5G Heterogeneous Networks
Andrabi U., Wadood E., Ojha S. K. et al., IEEE Access 2026 Vol. 14 P. 103358–103375
The emergence of 5G networks, aimed at accommodating diverse service requirements such as enhanced Mobile Broadband (eMBB), Ultra-Reliable Low-Latency Communication (URLLC), and massive Machine-Type Communication (mMTC), has presented significant challenges in radio resource management and network slicing. In dynamic heterogeneous network systems, traditional heuristics and mathematical programming methods find it challenging to attain scalable multi-objective ...
Added: September 23, 2026
LLM-assisted writing and citation advantage: evidence from scientific publications before and after ChatGPT release
Paklina S., Parshakov P., Elena Rapoport, Scientometrics 2026 P. 1–26
Generative artificial intelligence has become a routine part of academic writing. While much of the debate has focused on questions of integrity and authorship, less attention has been paid to how AI-assisted writing may affect research evaluation itself. This paper asks a straightforward but important question: does the use of LLMs in academic writing change ...
Added: September 23, 2026
Label-Free Quantification in the Crux Toolkit
Kertesz-Farkas A., Acquaye F. L., Journal of Proteome Research 2026 Vol. 25 P. 3764–3768
Ultimately, most tandem mass spectrometry (MS/MS) proteomics experiments aim to not just detect but also quantify the proteins in a given complex sample. Here, we describe an extension to the Crux MS/MS analysis toolkit to enable label-free quantification of peptides. We demonstrate that Crux’s new quantification command, which is modeled after the algorithms implemented in ...
Added: September 23, 2026
Risk Assessment Models for Heated Tobacco Products
Maddalena L., Yildiz B., Del Vecchio Blanco F. et al., Risk Analysis 2026 Vol. 46 No. 4 P. 1–26
Heated tobacco products (HTPs) are marketed as alternatives to conventional cigarettes with a potential reduced risk profile. Yet, their actual impact on cancer and noncancer disease risk remains uncertain and requires rigorous quantitative assessment. In this study, we develop a unified and transparent computational framework for toxicological risk assessment of HTPs, integrating chemical emissions data ...
Added: September 22, 2026
Обобщение пространства Фока
Дильмухаметова Алия Мидхатовна, Напалков В. В., Муллабаева А. У., Уфимский математический журнал 2010 Т. 2 № 1 С. 52–58
В данной статье введены обобщённые пространства Фока и рассмотрены основные свойства этих пространств. Найдена операция, сопряженная к операции умножения на переменную в обобщенном пространстве Фока. Также определены собственные функции сопряженного оператора. Изучены обобщенное преобразование Лапласа и задача построения базиса для введенных пространств. ...
Added: September 21, 2026
Segmentation of the Iris and Pupil of the Human Eye in Images from an Infrared Camera
Aleksei Samarin, Alexander Savelev, Aleksei Toropov et al., Pattern Recognition and Image Analysis 2024 Vol. 34 No. 3 P. 855–862
Tasks related to the automation of medical data processing are becoming more urgent. Particular attention is paid to systems for monitoring and analyzing human physiological parameters. Such systems often use specialized sensors to capture biomedical images, such as infrared cameras. This article describes our study of the problem of segmenting the eye pupil and iris ...
Added: September 21, 2026
A Model Based on Universal Filters for Image Color Correction
Aleksei Samarin, Nazarenko A., Alexander Savelev et al., Pattern Recognition and Image Analysis 2024 Vol. 34 No. 3 P. 844–854
Improving image quality is becoming an increasingly popular task, especially when working with mobile devices. One common approach to image enhancement is the use of convolutional neural networks. However, to achieve good results, such networks must be large enough, otherwise there is a risk of unwanted artifacts. In addition, large convolutional neural networks require significant ...
Added: September 21, 2026
Streptococci Recognition in Microscope Images Using Taxonomy-based Visual Features
Aleksei Samarin, Alexander Savelev, Aleksei Toropov et al., Optical Memory and Neural Networks (Information Optics) 2024 Vol. 33 P. 424–434
This study explores the development of classifiers for microbial images, specifically focusing on streptococci captured via microscopy of live samples. Our approach uses AutoML-based techniques and automates the creation and analysis of feature spaces to produce optimal descriptors for classifying these microscopic images. This technique leverages interpretable taxonomic features based on the external geometric attributes ...
Added: September 21, 2026
Specialized Image Descriptors Adaptation for Polyp Recognition over Endoscopic Images
Aleksei Samarin, Aleksei Toropov, Alexander Savelev et al., Pattern Recognition and Image Analysis 2024 Vol. 34 No. 4 P. 1053–1060
This paper presents a novel approach to classification in biomedical imaging, specifically targeting polyp recognition in video endoscopy snapshots. Our method leverages specialized image descriptors to enhance the accuracy and robustness of polyp recognition. By employing these specialized descriptors, we address the challenges inherent in analyzing biomedical images from open datasets. Our approach not only ...
Added: September 21, 2026
Разработка микросервиса ADP для идентификации источников выбросов на основе машинного обучения с подкреплением
Kychkin A., Chernitsin I., Прикладная информатика 2026 № 1(121) С. 40–58
The results of the development of a software microservice embedded in atmospheric air quality monitoring systems to support the identification of industrial pollution sources are presented. The emission and subsequent spread of harmful substances in the lower layers of the atmosphere is dynamic and characterized by high uncertainty due to the specific features of technological ...
Added: April 23, 2026
Artificial Neural Networks and Machine Learning. ICANN 2025 International Workshops and Special Sessions: 34th International Conference on Artificial Neural Networks, Kaunas, Lithuania, September 9–12, 2025, Proceedings, Part V
Cham: Springer, 2025.
This book constitutes the refereed proceedings of 34th International Workshops which were held in conjunction with the 34th International Conference on Artificial Neural Networks and Machine Learning, ICANN 2025, held in Kaunas, Lithuania, September 9–12, 2025.   The 20 full papers and 8 abstracts included in this workshop volume were carefully reviewed and selected from 42 submissions. ...
Added: September 29, 2025
Analysis of a Company Model in Conditions of Unstable Demand Using Reinforcement Learning Methods
Delev A., Semakov S., , in: 2025 8th International Conference on Artificial Intelligence and Big Data (ICAIBD).: IEEE, 2025. P. 318–322.
Profit is one of the most important economic indicators of a company’s performance, and for every company it is necessary to allocate resources in such a way as to obtain the maximum possible profit. The profit maximization problem is usually a dynamic optimization problem. This article discusses an approach to solving the production expansion problem ...
Added: August 25, 2025
Pseudo-collusion in a centralized algorithmic financial market
Pastushkov A., Boulatov A., Finance Research Letters 2025 Vol. 83 Article 107671
Recent studies have increasingly explored whether reinforcement learning algorithms can give rise to cooperative behavior that results in non-competitive pricing across various market settings. In financial markets, Cartea et al. (2022) show that market makers using multi-armed bandit (MAB) algorithms generally converge to competitive pricing in quote-driven over-the-counter (OTC) markets, barring some unlikely exceptions where ...
Added: June 19, 2025
The beer game bullwhip effect mitigation: a deep reinforcement learning approach
Rozhkov M., Alyamovskaya N., Zakhodiakin G., International Journal of Production Research 2025 Vol. 63 No. 18 P. 6630–6647
This article investigates the application of reinforcement learning (RL) methods to optimise a four-echelon linear supply chain model with stochastic demand. The proposed supply chain configuration is largely based on the production-distribution supply chain of the MIT Supply Chain Beer Game. We show that RL can significantly improve ordering efficiency and overall supply chain performance. ...
Added: March 24, 2025
Deep Reinforcement Learning-Based Congestion Control for File Transfer over QUIC
Blokhin A., Kalev V., Pusev R. et al., , in: 2024 IEEE International Multi-Conference on Engineering, Computer and Information Sciences (SIBIRCON).: Novosibirsk: IEEE, 2024. P. 25–30.
Congestion control is one of the key mechanisms of communication in QUIC protocol which controls how much data and at which rate can be send to an endpoint at particular moment of time for better use of shared network resources and avoids moving into congestive collapse state. In this work we tackle the problem of ...
Added: December 18, 2024
Generative Flow Networks as Entropy-Regularized RL
Tiapkin D., Morozov N., Naumov A. et al., , in: Proceedings of The 27th International Conference on Artificial Intelligence and Statistics (AISTATS 2024), 2-4 May 2024, Palau de Congressos, Valencia, Spain. PMLR: Volume 238Vol. 238.: Valencia: PMLR, 2024. P. 4213–4221.
The recently proposed generative flow networks (GFlowNets) are a method of training a policy to sample compositional discrete objects with probabilities proportional to a given reward via a sequence of actions. GFlowNets exploit the sequential nature of the problem, drawing parallels with reinforcement learning (RL). Our work extends the connection between RL and GFlowNets to ...
Added: June 22, 2024
Model-free Posterior Sampling via Learning Rate Randomization
Tiapkin D., Belomestny D., Calandriello D. et al., , in: Advances in Neural Information Processing Systems 36 (NeurIPS 2023).: Curran Associates, Inc., 2023. P. 73719–73774.
Added: February 17, 2024
Randomized Machine Learning Algorithms to Forecast the Evolution of Thermokarst Lakes Area in Permafrost Zones
Yu. A. Dubnov, A. Yu. Popkov, Polishchuk V. Y. et al., Automation and Remote Control 2023 Vol. 84 No. 1 P. 64–81
Randomized machine learning focuses on problems with considerable uncertainty in data and models. Machine learning algorithms are formulated in terms of a functional entropylinear programming problem. We adapt these algorithms to forecasting problems on an example of the evolution of thermokarst lakes area in permafrost zones. Thermokarst lakes generate methane, a greenhouse gas affecting climate ...
Added: February 5, 2024
Fast Rates for Maximum Entropy Exploration
Tiapkin D., Belomestny D., Calandriello D. et al., , in: Proceedings of the 40th International Conference on Machine Learning: Volume 202: International Conference on Machine Learning, 23-29 July 2023, Honolulu, Hawaii, USAVol. 202: International Conference on Machine Learning, 23-29 July 2023, Honolulu, Hawaii, USA.: PMLR, 2023. P. 34161–34221.
Added: December 1, 2023
Sharp Deviations Bounds for Dirichlet Weighted Sums with Application to analysis of Bayesian algorithms
Tiapkin D., Belomestny D., Naumov A. et al., Working papers by Cornell University. Series math "arxiv.org" 2023 Article 2304.03056
In this work, we derive sharp non-asymptotic deviation bounds for weighted sums of Dirichlet random variables. These bounds are based on a novel integral representation of the density of a weighted Dirichlet sum. This representation allows us to obtain a Gaussian-like approximation for the sum distribution using geometry and complex analysis methods. Our results generalize ...
Added: June 28, 2023
  • About
  • About
  • Key Figures & Facts
  • Sustainability at HSE University
  • Faculties & Departments
  • International Partnerships
  • Faculty & Staff
  • HSE Buildings
  • HSE University for Persons with Disabilities
  • Public Enquiries
  • Studies
  • Admissions
  • Programme Catalogue
  • Undergraduate
  • Graduate
  • Exchange Programmes
  • Summer University
  • Summer Schools
  • Semester in Moscow
  • Business Internship
  • Research
  • International Laboratories
  • Research Centres
  • Research Projects
  • Monitoring Studies
  • Conferences & Seminars
  • Academic Jobs
  • Yasin (April) International Academic Conference on Economic and Social Development
  • Media & Resources
  • Publications by staff
  • HSE Journals
  • Publishing House
  • iq.hse.ru: commentary by HSE experts
  • Library
  • Economic & Social Data Archive
  • Video
  • HSE Repository of Socio-Economic Information
  • HSE1993–2026
  • Contacts
  • Copyright
  • Privacy Policy
  • Site Map
Edit