Deep Reinforcement Learning-Based Congestion Control for File Transfer over QUIC

Blokhin A.; Kalev V.; R. Pusev; Kviatkovsky I.; Moskvitin D.

doi:10.1109/SIBIRCON63777.2024.10758499

Publications

?

Deep Reinforcement Learning-Based Congestion Control for File Transfer over QUIC

P. 25–30.

Blokhin A., Kalev V., Pusev R., Kviatkovsky I., Moskvitin D.

Congestion control is one of the key mechanisms of communication in QUIC protocol which controls how much data and at which rate can be send to an endpoint at particular moment of time for better use of shared network resources and avoids moving into congestive collapse state. In this work we tackle the problem of congestion control for file transfer over QUIC. We propose Reinforcement Learning Soft Actor Critique based congestion control with monitor window limit (SAC-MWL) and address the challenges it may have in a real network. We show how these challenges can block RL-based congestion control from effective learning of best policy and suggested a way to solve it. Our experiments are conducted in three different domains: pure virtual environment, lab-controlled network and real network where end points are spread all over the world. We compared performance of our approach with classical congestion controls CUBIC and BBR in a various network conditions and achieved up to 66.9 % reduction in a file transfer time.

Keywords: networking reinforcement learning QUIC congestion control file transfer mininet

In book

2024 IEEE International Multi-Conference on Engineering, Computer and Information Sciences (SIBIRCON)

Novosibirsk: IEEE, 2024.

Разработка микросервиса ADP для идентификации источников выбросов на основе машинного обучения с подкреплением

Kychkin A., Chernitsin I., Прикладная информатика 2026 № 1(121) С. 40–58

The results of the development of a software microservice embedded in atmospheric air quality monitoring systems to support the identification of industrial pollution sources are presented. The emission and subsequent spread of harmful substances in the lower layers of the atmosphere is dynamic and characterized by high uncertainty due to the specific features of technological ...

Added: April 23, 2026

Methodology for Experimental Comparison of Redundancy for QUIC and TCP Protocols

Burkov Artem A., Rogov Danil V., Sharipov Kamil A. et al., , in: 2025 XIХ International Symposium on Problems of Redundancy in Information and Control Systems (Redundancy), 5-7 Nov. 2025.: IEEE, 2025. P. 1–7.

The paper presents a comparison of the redundancy introduced when transmitting data using the QUIC and TCP protocols. Issues concerning the impact of redundancy—including that introduced by error-corretcing codes—on protocol efficiency are discussed. To quantitatively assess protocol performance, the protocol efficiency coefficient is introduced, defined as the ratio of the volume of payload data to ...

Added: December 1, 2025

Digital strategic collaborations in agriculture: a novel asset for local identity enhancement toward Agrifood 5.0

Cuomo M. T., Genovino C., De Andreis F. et al., British Food Journal 2024 Vol. 126 No. 11 P. 3922–3952

Purpose The aim of this research is to elucidate the correlation between open innovation, digital strategies and networking in enhancing agricultural enterprises within the new perspective of Agrifood 5.0. As such, it contributes to making businesses more competitive, especially in the Italian agricultural sector, where small and medium-sized enterprises are highly fragmented. Numerous studies have asserted ...

Added: October 2, 2025

Artificial Neural Networks and Machine Learning. ICANN 2025 International Workshops and Special Sessions: 34th International Conference on Artificial Neural Networks, Kaunas, Lithuania, September 9–12, 2025, Proceedings, Part V

Cham: Springer, 2025.

This book constitutes the refereed proceedings of 34th International Workshops which were held in conjunction with the 34th International Conference on Artificial Neural Networks and Machine Learning, ICANN 2025, held in Kaunas, Lithuania, September 9–12, 2025. The 20 full papers and 8 abstracts included in this workshop volume were carefully reviewed and selected from 42 submissions. ...

Added: September 29, 2025

Analysis of a Company Model in Conditions of Unstable Demand Using Reinforcement Learning Methods

Delev A., Semakov S., , in: 2025 8th International Conference on Artificial Intelligence and Big Data (ICAIBD).: IEEE, 2025. P. 318–322.

Profit is one of the most important economic indicators of a company’s performance, and for every company it is necessary to allocate resources in such a way as to obtain the maximum possible profit. The profit maximization problem is usually a dynamic optimization problem. This article discusses an approach to solving the production expansion problem ...

Added: August 25, 2025

Pseudo-collusion in a centralized algorithmic financial market

Pastushkov A., Boulatov A., Finance Research Letters 2025 Vol. 83 Article 107671

Recent studies have increasingly explored whether reinforcement learning algorithms can give rise to cooperative behavior that results in non-competitive pricing across various market settings. In financial markets, Cartea et al. (2022) show that market makers using multi-armed bandit (MAB) algorithms generally converge to competitive pricing in quote-driven over-the-counter (OTC) markets, barring some unlikely exceptions where ...

Added: June 19, 2025

Reliable Queuing One-Way Delay Metric for Computer Networks

Kulya M., Pusev R., Moskvitin D., , in: 2025 International Russian Smart Industry Conference (SmartIndustryCon).: Sochi: IEEE, 2025. P. 83–88.

We claim the method to obtain a reliable queuing one-way delay metric for the wireless computer networking systems. The method requires only a measurement of the timestamp series between packet departure and arrival events. These timestamps are included inside the packets with stream frames. Thus, the method does not need to involve any additional probe ...

Added: May 19, 2025

The beer game bullwhip effect mitigation: a deep reinforcement learning approach

Rozhkov M., Alyamovskaya N., Zakhodiakin G., International Journal of Production Research 2025 Vol. 63 No. 18 P. 6630–6647

This article investigates the application of reinforcement learning (RL) methods to optimise a four-echelon linear supply chain model with stochastic demand. The proposed supply chain configuration is largely based on the production-distribution supply chain of the MIT Supply Chain Beer Game. We show that RL can significantly improve ordering efficiency and overall supply chain performance. ...

Added: March 24, 2025

Generative Flow Networks as Entropy-Regularized RL

Tiapkin D., Morozov N., Naumov A. et al., , in: Proceedings of The 27th International Conference on Artificial Intelligence and Statistics (AISTATS 2024), 2-4 May 2024, Palau de Congressos, Valencia, Spain. PMLR: Volume 238Vol. 238.: Valencia: PMLR, 2024. P. 4213–4221.

The recently proposed generative flow networks (GFlowNets) are a method of training a policy to sample compositional discrete objects with probabilities proportional to a given reward via a sequence of actions. GFlowNets exploit the sequential nature of the problem, drawing parallels with reinforcement learning (RL). Our work extends the connection between RL and GFlowNets to ...

Added: June 22, 2024

Model-free Posterior Sampling via Learning Rate Randomization

Tiapkin D., Belomestny D., Calandriello D. et al., , in: Advances in Neural Information Processing Systems 36 (NeurIPS 2023).: Curran Associates, Inc., 2023. P. 73719–73774.

Added: February 17, 2024

Reinforcement Procedure for Randomized Machine Learning

Yuri S. Popkov, Dubnov Y. A., Alexey Yu. Popkov, Mathematics 2023 Vol. 11 No. 17 Article 3651

This paper is devoted to problem-oriented reinforcement methods for the numerical implementation of Randomized Machine Learning. We have developed a scheme of the reinforcement procedure based on the agent approach and Bellman’s optimality principle. This procedure ensures strictly monotonic properties of a sequence of local records in the iterative computational procedure of the learning process. ...

Added: February 5, 2024

Fast Rates for Maximum Entropy Exploration

Tiapkin D., Belomestny D., Calandriello D. et al., , in: Proceedings of the 40th International Conference on Machine Learning: Volume 202: International Conference on Machine Learning, 23-29 July 2023, Honolulu, Hawaii, USAVol. 202: International Conference on Machine Learning, 23-29 July 2023, Honolulu, Hawaii, USA.: PMLR, 2023. P. 34161–34221.

Added: December 1, 2023

Sharp Deviations Bounds for Dirichlet Weighted Sums with Application to analysis of Bayesian algorithms

Tiapkin D., Belomestny D., Naumov A. et al., Working papers by Cornell University. Series math "arxiv.org" 2023 Article 2304.03056

In this work, we derive sharp non-asymptotic deviation bounds for weighted sums of Dirichlet random variables. These bounds are based on a novel integral representation of the density of a weighted Dirichlet sum. This representation allows us to obtain a Gaussian-like approximation for the sum distribution using geometry and complex analysis methods. Our results generalize ...

Added: June 28, 2023

Variance Reduction for Policy-Gradient Methods via Empirical Variance Minimization

Belomestny D., Kaledin M., Golubev A., /. 2022.

Policy-gradient methods in Reinforcement Learning(RL) are very universal and widely applied in practice but their performance suffers from the high variance of the gradient estimate. Several procedures were proposed to reduce it including actor-critic(AC) and advantage actor-critic(A2C) methods. Recently the approaches have got new perspective due to the introduction of Deep RL: both new control ...

Added: April 14, 2023

A note on observational equivalence of micro assumptions on macro level

Ponomarenko A. A., Economics: The Open-Access, Open-Assessment E-Journal 2020 Vol. 14 P. 1–15

The author set up a simplistic agent-based model where agents learn with reinforcement observing an incomplete set of variables. The model is employed to generate an artificial dataset that is used to estimate standard macro econometric models. The author shows that the results are qualitatively indistinguishable (in terms of the signs and significances of the ...

Added: March 28, 2023

Ambiguous tDCS: variability of the transcranial direct current stimulation effects in a reinforcement learning task

Anastasia Grigoreva, Aleksei Gorin, Valeriy Klyuchnikov et al., Brain Stimulation 2023 Vol. 16 No. 1 P. 273

Transcranial electrical stimulation (TES) is a popular approach for studying and modulating cortical function. According to somatic doctrine, anodal TES increases, while cathodal reduces cortical excitability. Currently, numerous studies use TES in behavioral experiments with no physiological control, relying on the assumption of fairness and complete predictability of stimulation models. However, control reveals the actual ...

Added: March 1, 2023