• A
  • A
  • A
  • АБВ
  • АБВ
  • АБВ
  • A
  • A
  • A
  • A
  • A
Обычная версия сайта
  • RU
  • EN
  • HSE University
  • Publications
  • Book chapter
  • Learning to Play Pong Video Game via Deep Reinforcement Learning: Tweaking Deep Q-Networks versus Episodic Control
  • RU
  • EN
Расширенный поиск
Высшая школа экономики
Национальный исследовательский университет
Priority areas
  • business informatics
  • economics
  • engineering science
  • humanitarian
  • IT and mathematics
  • law
  • management
  • mathematics
  • sociology
  • state and public administration
by year
  • 2028
  • 2027
  • 2026
  • 2025
  • 2024
  • 2023
  • 2022
  • 2021
  • 2020
  • 2019
  • 2018
  • 2017
  • 2016
  • 2015
  • 2014
  • 2013
  • 2012
  • 2011
  • 2010
  • 2009
  • 2008
  • 2007
  • 2006
  • 2005
  • 2004
  • 2003
  • 2002
  • 2001
  • 2000
  • 1999
  • 1998
  • 1997
  • 1996
  • 1995
  • 1994
  • 1993
  • 1992
  • 1991
  • 1990
  • 1989
  • 1988
  • 1987
  • 1986
  • 1985
  • 1984
  • 1983
  • 1982
  • 1981
  • 1980
  • 1979
  • 1978
  • 1977
  • 1976
  • 1975
  • 1974
  • 1973
  • 1972
  • 1971
  • 1970
  • 1969
  • 1968
  • 1967
  • 1966
  • 1965
  • 1964
  • 1963
  • 1958
  • More
Subject
News
September 30, 2026
'We Did Not Limit the Time for Questions'
The International Laboratory for Supercomputer Atomistic Modelling and Multi-Scale Analysis at HSE University held a major conference on molecular dynamics. Participants had the opportunity to attend all the presentations, while speakers were given as much time as they needed to answer questions. The HSE News Service interviewed Grigory Smirnov, Head of the Laboratory, and Genri Norman, Chief Research Fellow, about the conference preparations and the discussions it generated.
September 25, 2026
AI Users Earn Up to 41.8% More Than Non-Users
Research conducted by economists at HSE University has revealed a significant correlation between the regular use of GenAI in the workplace and higher pay among Russian employees. The study found that individuals who frequently use GenAI in their professional activities earn notably more than those who reject these new tools or resort to them occasionally. The salary premium for highly qualified specialists reaches 41.8%. The article was published in the Voprosy Ekonomiki journal.
September 24, 2026
‘Feedback and Constructive Criticism Are Essential in Our Profession
Vincent Fardeau, Associate Professor at HSE ICEF, has reached a major career milestone: he recently published his paper ‘Asymmetric Thin Markets’ in the Journal of Financial Economics, successfully passed his major academic review, and received tenure. In this interview, Vincent discusses the story behind the paper, explains the concept of asymmetric thin markets, and shares his advice for young scholars aiming to publish in top-tier journals.

 

Have you spotted a typo?
Highlight it, click Ctrl+Enter and send us a message. Thank you for your help!

Publications
  • Books
  • Articles
  • Chapters of books
  • Working papers
  • Report a publication
  • Research at HSE

?

Learning to Play Pong Video Game via Deep Reinforcement Learning: Tweaking Deep Q-Networks versus Episodic Control

P. 236–241.
Makarov I., Andrej Kashin, Alice Korinevskaya

We consider deep reinforcement learning algorithms for playing a game based on video input. We compare choosing proper hyper-parameters in deep Q-network model and model-free episodic control focused on reusing of successful strategies. The evaluation was made based on Pong video game implemented in Unreal Engine 4.

Language: English
Full text
Text on another site
Keywords: Unreal engine 4Deep Reinforcement LearningDeep Q-NetworksQ-LearningEpisodic ControlPong Video Game

In book

Supplementary Proceedings of the Sixth International Conference on Analysis of Images, Social Networks and Texts (AIST-SUP 2017), Moscow, Russia, July 27-29, 2017
Vol. 1975. , Aachen: CEUR-WS.org, 2017.
Similar publications
RL-ABC: Reinforcement learning for accelerator beamline control
Ibrahim A., Derkach D., Kaledin M. et al., Computer Physics Communications 2026 Vol. 329 Article 110407
We present RLABC (Reinforcement Learning for Accelerator Beamline Control), an open-source Python framework that automatically transforms standard Elegant beamline configurations into reinforcement learning environments. RLABC integrates with the widely-used Elegant beam dynamics simulation code via SDDS-based interfaces, enabling researchers to apply modern RL algorithms to beamline optimization with minimal RL-specific development. Validation on a test ...
Added: September 20, 2026
Обзор выпуклой оптимизации марковских процессов принятия решений
Rudenko V., Yudin N., Васин А. А., Компьютерные исследования и моделирование 2023 Т. 15 № 2 С. 329–353
This article reviews both historical achievements and modern results in the field of Markov Decision Process (MDP) and convex optimization. This review is the first attempt to cover the field of reinforcement learning in Russian in the context of convex optimization. The fundamental Bellman equation and the criteria of optimality of policy — strategies based on it, ...
Added: November 29, 2024
MineRL Diamond 2021 Competition: Overview, Results, and Lessons Learned
Nikulin A. M., Belousov Y., Svidchenko O. et al., , in: Proceedings of the NeurIPS 2021 Competitions and Demonstrations Track.: PMLR, 2022.
Reinforcement learning competitions advance the field by providing appropriate scope and support to develop solutions toward a specific problem. To promote the development of more broadly applicable methods, organizers need to enforce the use of general techniques, the use of sample-efficient methods, and the reproducibility of the results. While beneficial for the research community, these ...
Added: October 8, 2024
Deep Reinforcement Learning with DQN vs. PPO in VizDoom
Anton Zakharenkov, Makarov I., , in: Proceedings of IEEE 21st International Symposium on Computational Intelligence and Informatics (CINTI'21), 18-20 Nov. 2021.: NY: IEEE, 2021. P. 000131–000136.
Added: January 19, 2022
Flatland Competition 2020: MAPF and MARL for Efficient Train Coordination on a Grid World
Laurent F., Schneider M., Scheller C. et al., , in: Proceedings of Machine Learning ResearchVol. 133: Proceedings of the NeurIPS 2020: Competition and Demonstration Track.: PMLR, 2021. P. 275–301.
The Flatland competition aimed at finding novel approaches to solve the vehicle re-scheduling problem (VRSP). The VRSP is concerned with scheduling trips in traffic networks and the re-scheduling of vehicles when disruptions occur, for example the breakdown of a vehicle. While solving the VRSP in various settings has been an active area in operations research ...
Added: September 6, 2021
Deep Reinforcement Learning in VizDoom via DQN and Actor-Critic Agents
Maria Bakhanova, Ilya Makarov, , in: Advances in Computational Intelligence: 16th International Work-Conference on Artificial Neural Networks, IWANN 2021, Virtual Event, June 16–18, 2021, Proceedings, Part I* 1. Vol. 12861.: Springer, 2021. Ch. 12 P. 138–150.
In this work, we study the problem of learning reinforcement learning-based agents in a first-person shooter environment VizDoom. We compare several well-known architectures, such as DQN, DDQN, A3C, and Curiosity-driven model, while highlighting the main differences in learned policies of agents trained via these models. ...
Added: September 1, 2021
Balancing Rational and Other-Regarding Preferences in Cooperative-Competitive Environments
Ivanov D., Egorov V., Shpilman A., , in: AAMAS'2021: Proceedings of the 20th International Conference on Autonomous Agents and MultiAgent Systems.: IFAAMAS, 2021. P. 1536–1538.
Recent reinforcement learning studies extensively explore the interplay between cooperative and competitive behaviour in mixed environments. Unlike cooperative environments where agents strive towards a common goal, mixed environments are notorious for the conflicts of selfish and social interests. As a consequence, purely rational agents often struggle to maintain cooperation. A prevalent approach to induce cooperative ...
Added: May 29, 2021
AAMAS'2021: Proceedings of the 20th International Conference on Autonomous Agents and MultiAgent Systems
IFAAMAS, 2021.
These are the proceedings of the 20th International Conference on Autonomous Agents and Multiagent Systems (AAMAS-2021). They are published by the International Foundation for Autonomous Agents and Multiagent Systems (IFAAMAS). ...
Added: May 29, 2021
Workshop on AI for Autonomous Driving (AIAD)
[б.и.], 2020.
Self-driving cars and advanced safety features present one of today’s greatest challenges and opportunities for Artificial Intelligence (AI). Despite billions of dollars of investments and encouraging progress under certain operational constraints, there are no driverless cars on public roads today without human safety drivers. Autonomous Driving research spans a wide spectrum, from modular architectures -- ...
Added: December 28, 2020
MAGNet: Multi-Agent Graph Network for Deep Multi-Agent Reinforcement Learning
Shpilman A., Malysheva A., Kudenko D., , in: Proceedings of 2019 XVI International Symposium "Problems of Redundancy in Information and Control Systems" (REDUNDANCY).: IEEE, 2019. P. 171–176.
Over recent years, deep reinforcement learning has shown strong successes in complex single-Agent tasks, and more recently this approach has also been applied to multi-Agent domains. In this paper, we propose a novel approach, called MAGNet, to multi-Agent reinforcement learning that utilizes a relevance graph representation of the environment obtained by a self-Attention mechanism, and ...
Added: July 15, 2020
Deep Reinforcement Learning Methods in Match-3 Game
Ildar Kamaldinov, Makarov I., , in: Analysis of Images, Social Networks and Texts. 8th International Conference AIST 2019.: Springer, 2019. P. 51–62.
A large number of methods are being developed in the deep reinforcement learning area recently, but the scope of their application is limited. The number of environments does not always allow for a comprehensive assessment of a new agent training algorithm. The main purpose of this article is to present another environment for Match-3 game ...
Added: February 4, 2020
Artificial Intelligence for Prosthetics: Challenge Solutions
Shpilman A., Kidzinski L., Ong C. et al., , in: The NeurIPS '18 Competition: From Machine Learning to Intelligent Conversations.: Springer, 2020. P. 69–128.
Added: December 2, 2019
Deep Reinforcement Learning with VizDoom First-Person Shooter
Dmitry Akimov, Makarov I., , in: Proceedings of the Fifth Workshop on Experimental Economics and Machine Learning at the National Research University Higher School of Economics co-located with the Seventh International Conference on Applied Research in Economics (iCare7).: Aachen: CEUR Workshop Proceedings, 2019. P. 3–17.
In this work, we study deep reinforcement algorithms for partially observable Markov decision processes (POMDP) combined with Deep Q-Networks. To our knowledge, we are the first to apply standard Markov decision process architectures to POMDP scenarios. We propose an extension of DQN with Dueling Networks and several other model-free policies to training agent using deep ...
Added: November 19, 2019
Deep Reinforcement Learning in Match-3 Game
Ildar Kamaldinov, Makarov I., , in: Procedings of IEEE Conference on Games (COG'19).: NY: IEEE, 2019. P. 1–4.
An increasing number of algorithms in deep reinforcement learning area creates new challenges for environments, particularly, for their comprehensive analysis and searching application areas. The key purpose of this article is to provide an extensible environment for researches. We consider a Match-3 game, which has simple gameplay, but challenging game design for engaging players. The ...
Added: July 30, 2019
Deep Reinforcement Learning in VizDoom First-Person Shooter for Health Gathering Scenario
Dmitry Akimov, Makarov I., , in: Proceedings of 11th International Conference on Advances in Multimedia (MMEDIA'19).: Lansing: ThinkMind, 2019. P. 59–64.
In this work, we study the effect of combining existent improvements for Deep Q-Networks (DQN) in Markov Decision Processes (MDP) and Partially Observable MDP (POMDP) settings. Combinations of several heuristics, such as Distributional Learning and Dueling architectures improvements, for MDP are well-studied. We propose a new combination method of simple DQN extensions and develop a ...
Added: July 29, 2019
MAGNet: Multi-agent Graph Network for Deep Multi-agent Reinforcement Learning
Shpilman A., Malysheva A., Kudenko D., , in: Adaptive and Learning Agents Workshop at International Joint Conference on Autonomous Agents and Multiagent Systems.: [б.и.], 2019. P. 1–8.
Over recent years, deep reinforcement learning has shown strong successes in complex single-agent tasks, and more recently this approach has also been applied to multi-agent domains. In this paper, we propose a novel approach, called MAGNet, to multi-agent reinforcement learning that utilizes a relevance graph representation of the environment obtained by a self-attention mechanism, and ...
Added: June 13, 2019
Unreal Engine VR для разработчиков
Макеффри М., М.: Эксмо, 2019.
Первое руководство по созданию виртуальной реальности с использованием движка Unreal Engine 4 на русском языке! VR - новый, удивительный рубеж для разработчиков игр и специалистов по визуализации. А Unreal Engine 4 - идеальная платформа для этого. "Unreal Engine VR. Руководство по разработке" - это исчерпывающее руководство по созданию потрясающих приложений на любых VR-устройствах, совместимых с Unreal ...
Added: May 16, 2019
Deep Multi-Agent Reinforcement Learning with Relevance Graphs
Shpilman A., Malysheva A., Sung T. T. et al., , in: Deep RL Workshop NeurIPS 2018.: [б.и.], 2018. P. 1–10.
Over recent years, deep reinforcement learning has shown strong successes in complex single-agent tasks, and more recently this approach has also been applied to multi-agent domains. In this paper, we propose a novel approach, called MAGnet, to multi-agent reinforcement learning (MARL) that utilizes a relevance graph representation of the environment obtained by a self-attention mechanism, and a message-generation technique inspired ...
Added: January 18, 2019
Learning to Run with Reward Shaping from Video Data
Malysheva A., Shpilman A., Kudenko D., , in: ALA 2018 - Workshop at the Federated AI Meeting 2018.: ALA, 2018. P. 1–7.
Learning to produce efficient movement behaviour for humanoid robots from scratch is a hard problem, as has been illustrated by the "Learning to run" competition at NIPS 2017. The goal of this competition was to train a two-legged model of a humanoid body to run in a simulated race course with maximum speed. All submissions ...
Added: October 16, 2018
Adapting First-Person Shooter Video Game for Playing with Virtual Reality Headsets
Makarov I., Konoplya O., Pavel Polyakov et al., , in: Proceedings of the Thirtieth International Florida Artificial Intelligence Research Society Conference, FLAIRS 2017, Marco Island, Florida, USA, May 22-24, 2017. AAAI Press 2017, ISBN 978-1-57735-787-2.: Palo Alto: AAAI Press, 2017. P. 412–415.
In this article a combination of two modern aspects of games development is considered: (i) the impact of high quality graphics and virtual reality (VR) user adaptation to believe in realness of in-game events by user’s own eyes; (ii) modeling an enemy’s behavior under automatic computer control, called BOT, which reacts similarly to human players. ...
Added: June 24, 2017
  • About
  • About
  • Key Figures & Facts
  • Sustainability at HSE University
  • Faculties & Departments
  • International Partnerships
  • Faculty & Staff
  • HSE Buildings
  • HSE University for Persons with Disabilities
  • Public Enquiries
  • Studies
  • Admissions
  • Programme Catalogue
  • Undergraduate
  • Graduate
  • Exchange Programmes
  • Summer University
  • Summer Schools
  • Semester in Moscow
  • Business Internship
  • Research
  • International Laboratories
  • Research Centres
  • Research Projects
  • Monitoring Studies
  • Conferences & Seminars
  • Academic Jobs
  • Yasin (April) International Academic Conference on Economic and Social Development
  • Media & Resources
  • Publications by staff
  • HSE Journals
  • Publishing House
  • iq.hse.ru: commentary by HSE experts
  • Library
  • Economic & Social Data Archive
  • Video
  • HSE Repository of Socio-Economic Information
  • HSE1993–2026
  • Contacts
  • Copyright
  • Privacy Policy
  • Site Map
Edit