• A
  • A
  • A
  • АБВ
  • АБВ
  • АБВ
  • A
  • A
  • A
  • A
  • A
Обычная версия сайта
  • RU
  • EN
  • HSE University
  • Publications
  • Book chapter
  • Improving GFlowNets with Monte Carlo Tree Search
  • RU
  • EN
Расширенный поиск
Высшая школа экономики
Национальный исследовательский университет
Priority areas
  • business informatics
  • economics
  • engineering science
  • humanitarian
  • IT and mathematics
  • law
  • management
  • mathematics
  • sociology
  • state and public administration
by year
  • 2028
  • 2027
  • 2026
  • 2025
  • 2024
  • 2023
  • 2022
  • 2021
  • 2020
  • 2019
  • 2018
  • 2017
  • 2016
  • 2015
  • 2014
  • 2013
  • 2012
  • 2011
  • 2010
  • 2009
  • 2008
  • 2007
  • 2006
  • 2005
  • 2004
  • 2003
  • 2002
  • 2001
  • 2000
  • 1999
  • 1998
  • 1997
  • 1996
  • 1995
  • 1994
  • 1993
  • 1992
  • 1991
  • 1990
  • 1989
  • 1988
  • 1987
  • 1986
  • 1985
  • 1984
  • 1983
  • 1982
  • 1981
  • 1980
  • 1979
  • 1978
  • 1977
  • 1976
  • 1975
  • 1974
  • 1973
  • 1972
  • 1971
  • 1970
  • 1969
  • 1968
  • 1967
  • 1966
  • 1965
  • 1964
  • 1963
  • 1958
  • More
Subject
News
September 25, 2026
AI Users Earn Up to 41.8% More Than Non-Users
Research conducted by economists at HSE University has revealed a significant correlation between the regular use of GenAI in the workplace and higher pay among Russian employees. The study found that individuals who frequently use GenAI in their professional activities earn notably more than those who reject these new tools or resort to them occasionally. The salary premium for highly qualified specialists reaches 41.8%. The article was published in the Voprosy Ekonomiki journal.
September 24, 2026
‘Feedback and Constructive Criticism Are Essential in Our Profession
Vincent Fardeau, Associate Professor at HSE ICEF, has reached a major career milestone: he recently published his paper ‘Asymmetric Thin Markets’ in the Journal of Financial Economics, successfully passed his major academic review, and received tenure. In this interview, Vincent discusses the story behind the paper, explains the concept of asymmetric thin markets, and shares his advice for young scholars aiming to publish in top-tier journals.
September 22, 2026
Personal Interest in Doctoral Thesis Topic Most Important for Confidence in Successful Defence
A researcher at HSE University analysed data on 1,539 doctoral students from 161 Russian universities to identify which features of a thesis topic are associated with academic success and engagement. The most important factor was found to be personal interest in the research topic, which was associated with almost all key aspects of doctoral programme experience—from engaging with the academic supervisor to research activity and confidence about successfully defending the thesis. The findings have been published in Higher Education.

 

Have you spotted a typo?
Highlight it, click Ctrl+Enter and send us a message. Thank you for your help!

Publications
  • Books
  • Articles
  • Chapters of books
  • Working papers
  • Report a publication
  • Research at HSE

?

Improving GFlowNets with Monte Carlo Tree Search

.
Morozov N., Tiapkin D., Samsonov S., Naumov A., Vetrov D.
Language: English
Text on another site
Keywords: generative modelingdeep reinforcement learningGenerative Flow Networks

In book

ICML 2024 Workshop on Structured Probabilistic Inference & Generative Modeling
OpenReview, 2024.
Similar publications
Learning-Based UAV–RIS Secure Communication Under Eavesdropper Location Uncertainty
Ehab S. Suleiman, Ali J. Dayoub, , in: Proceedings of the 2026 8th International Youth Conference on Radio Electronics, Electrical and Power Engineering (REEPE).: IEEE, 2026. Ch. 165 P. 1–6.
Unmanned aerial vehicle (UAV)-assisted reconfigurable intelligent surface (RIS) systems can enhance physical layer security through joint mobility and propagation control. However, most existing designs assume the availability of the eavesdropper's channel state information (CSI), which is unrealistic in passive eavesdropping scenarios. In this paper, secure UAV-RIS downlink communication is studied under bounded eavesdropper location uncertainty, ...
Added: April 30, 2026
Revisiting Non-Acyclic GFlowNets in Discrete Environments
Morozov N., Maximov I., Tiapkin D. et al., , in: Volume 267: International Conference on Machine Learning, 13-19 July 2025, Vancouver Convention Center, Vancouver, CanadaVol. 267.: [б.и.], 2025. P. 44887–44910.
Generative Flow Networks (GFlowNets) are a family of generative models that learn to sample objects from a given probability distribution, potentially known up to a normalizing constant. Instead of working in the object space, GFlowNets proceed by sampling trajectories in an appropriately constructed directed acyclic graph environment, greatly relying on the acyclicity of the graph. ...
Added: October 15, 2025
Optimizing Backward Policies in GFlowNets via Trajectory Likelihood Maximization
Timofei Gritsaev, Morozov N., Samsonov S. et al., , in: Proceedings of the 13th International Conference on Learning Representations (ICLR 2025).: ICLR, 2025. P. 95626–95646.
Generative Flow Networks (GFlowNets) are a family of generative models that learn to sample objects with probabilities proportional to a given reward function. The key concept behind GFlowNets is the use of two stochastic policies: a forward policy, which incrementally constructs compositional objects, and a backward policy, which sequentially deconstructs them. Recent results show a ...
Added: August 15, 2025
Optical stabilization for laser communication satellite systems through proportional–integral–derivative (PID) control and reinforcement learning approach
Бахшалиев Р. М., Reutov A., Vorobey S. et al., Review of Scientific Instruments 2025 Vol. 96 No. 3
One of the main issues of the satellite-to-ground optical communication, including free-space satellite quantum key distribution (QKD), is an achievement of the reasonable accuracy of positioning, navigation, and optical stabilization. Proportional–integral–derivative (PID) controllers can handle various control tasks in optical systems. Recent research shows the promising results in the area of composite control systems including ...
Added: May 13, 2025
Optimization of the Accelerator Control by Reinforcement Learning: A Simulation-Based Approach
Ibrahim A., Derkach D., Petrenko A. et al., Physics of Particles and Nuclei 2025 Vol. 56 No. 6 P. 1476–1481
Optimizing accelerator control is a critical challenge in experimental particle physics, requiring significant manual effort and resource expenditure. Traditional tuning methods are often time-consuming and reliant on expert input, highlighting the need for more efficient approaches. This study aims to create a simulation-based framework integrated with Reinforcement Learning (RL) to address these challenges. Using \texttt{Elegant} ...
Added: March 16, 2025
Generative models and seq2seq techniques for the flash-simulation of the LHCb experiment
Derkach D., Anderlini L., Capelli S. et al., Proceedings of Science 2025 Vol. 476 P. 1032
Simulating detector and reconstruction effects on physics quantities is crucial for data analysis, but it is coming unsustainably costly for the upcoming HEP experiments. The most radical approach to speed-up detector simulation is Flash Simulation, as proposed by the LHCb collaboration in Lamarr, a software package implementing a novel simulation paradigm relying on Deep Generative ...
Added: March 13, 2025
Adaptive Algorithm for Selecting the Optimal Trading Strategy Based on Reinforcement Learning for Managing a Hedge Fund
Belyakov B., Sizykh D., IEEE Access 2024 Vol. 12 P. 189047–189063
In hedge fund management, the ability to dynamically select optimal trading strategies is paramount for maximizing returns and mitigating risk. This paper presents a pioneering approach that integrates Reinforcement Learning (RL), specifically the Proximal Policy Optimization (PPO) algorithm, into the strategy selection process for hedge fund management. Our model considers a diverse array of strategies, ...
Added: January 15, 2025
The LHCb ultra-fast simulation option, Lamarr design and validation
Derkach D., Kazeev N., Mokhnenko S. et al., EPJ Web of Conferences 2024 Vol. 295 P. 03040
Detailed detector simulation is the major consumer of CPU resources at LHCb, having used more than 90% of the total computing budget during Run 2 of the Large Hadron Collider at CERN. As data is collected by the upgraded LHCb detector during Run 3 of the LHC, larger requests for simulated data samples are necessary, ...
Added: January 8, 2025
ICML 2024 Workshop on Structured Probabilistic Inference & Generative Modeling
OpenReview, 2024.
The workshop focuses on theory, methodology, and application of structured probabilistic inference and generative modeling. ...
Added: October 24, 2024
Star-Shaped Denoising Diffusion Probabilistic Models
Andrey Okhotin, Dmitry Molchanov, Arkhipkin V. et al., , in: Advances in Neural Information Processing Systems 36 (NeurIPS 2023).: Curran Associates, Inc., 2023. P. 10038–10067.
Added: February 15, 2024
When to Switch: Planning and Learning for Partially Observable Multi-Agent Pathfinding
Skrynnik A., Andreychuk A., Yakovlev K. et al., IEEE Transactions on Neural Networks and Learning Systems 2024 Vol. 35 No. 12 P. 17411–17424
Multi-agent pathfinding (MAPF) is a problem that involves finding a set of non-conflicting paths for a set of agents confined to a graph. In this work, we study a MAPF setting, where the environment is only partially observable for each agent, i.e., an agent observes the obstacles and other agents only within a limited field-of-view. ...
Added: December 4, 2023
Dealing With Sparse Rewards Using Graph Neural Networks
Gerasyov Matvey, Makarov I., IEEE Access 2023 Vol. 11 P. 89180–89187
Deep reinforcement learning in partially observable environments is a difficult task in itself and can be further complicated by a sparse reward signal. Most tasks involving navigation in three-dimensional environments provide the agent with minimal information. Typically, the agent receives a visual observation input from the environment and is rewarded once at the end of ...
Added: August 28, 2023
Artificial Intelligence and Mathematical Models of Power Grids Driven by Renewable Energy Sources: A Survey
Srinivasan S., Kumarasamy S., Andreadakis Z. et al., Energies 2023 Vol. 16 No. 14 Article 5383
To face the impact of climate change in all dimensions of our society in the near future, the European Union (EU) has established an ambitious target. Until 2050, the share of renewable power shall increase up to 75% of all power injected into nowadays’ power grids. While being clean and having become significantly cheaper, renewable ...
Added: July 17, 2023
Maximum Entropy Model-based Reinforcement Learning
Svidchenko O., Shpilman A., , in: NeurIPS'2021 Deep Reinforcement Learning Workshop.: [б.и.], 2021.
Recent advances in reinforcement learning have demonstrated its ability to solve hard agent-environment interaction tasks on a super-human level. However, the application of reinforcement learning methods to practical and real-world tasks is currently limited due to most RL state-of-art algorithms' sample inefficiency, i.e., the need for a vast number of training episodes. For example, OpenAI ...
Added: March 24, 2022
Self-Imitation Learning from Demonstrations
Ivanov D., Пшихачев Г. А., Егоров В. С. et al., , in: NeurIPS'2021 Deep Reinforcement Learning Workshop.: [б.и.], 2021.
Despite the numerous breakthroughs achieved with Reinforcement Learning (RL), solving environments with sparse rewards remains a challenging task that requires sophisticated exploration. Learning from Demonstrations (LfD) remedies this issue by guiding agent’s exploration towards states experienced by an expert. Naturally, the benefits of this approach hinge on the quality of demonstrations, which are rarely optimal ...
Added: March 24, 2022
NeurIPS'2021 Deep Reinforcement Learning Workshop
[б.и.], 2021.
Added: March 24, 2022
21st IEEE International Conference on Data Mining Workshops, ICDMW 2021
IEEE Computer Society, 2021.
The 21th IEEE International Conference on Data Mining (IEEE ICDM 2021) is a premier and truly international conference for researchers and practitioners in the broad area of data mining. The ICDM Workshops program (IEEE ICDMW) aims to provide a platform for multiple workshops with a range of more focused topics to be discussed and explored, where attendees can present ...
Added: February 4, 2022
Using RuGPT3-XL Model for RuNormAS competition
Emelyanov A., Shliazhko O., Katricheva N. et al., , in: Computational Linguistics and Intellectual Technologies: Papers from the Annual International Conference “Dialogue” (2021)Issue 20: Основной том.: -, 2021. Ch. 18 P. 204–212.
The paper presents a fine-tuning methodology of the RuGPT3-XL (Generative Pretrained Transformer-3 for Russian) language model for the normalization of text spans task. The solution is presented in a competition for two tasks: Normalization of Named Entities (Named entities) and Normalization of a wider class of text spans, including the normalization of different parts of ...
Added: September 5, 2021
  • About
  • About
  • Key Figures & Facts
  • Sustainability at HSE University
  • Faculties & Departments
  • International Partnerships
  • Faculty & Staff
  • HSE Buildings
  • HSE University for Persons with Disabilities
  • Public Enquiries
  • Studies
  • Admissions
  • Programme Catalogue
  • Undergraduate
  • Graduate
  • Exchange Programmes
  • Summer University
  • Summer Schools
  • Semester in Moscow
  • Business Internship
  • Research
  • International Laboratories
  • Research Centres
  • Research Projects
  • Monitoring Studies
  • Conferences & Seminars
  • Academic Jobs
  • Yasin (April) International Academic Conference on Economic and Social Development
  • Media & Resources
  • Publications by staff
  • HSE Journals
  • Publishing House
  • iq.hse.ru: commentary by HSE experts
  • Library
  • Economic & Social Data Archive
  • Video
  • HSE Repository of Socio-Economic Information
  • HSE1993–2026
  • Contacts
  • Copyright
  • Privacy Policy
  • Site Map
Edit