?
A Two-Stage Deep Reinforcement Learning Framework for Radio Resource Management and Network Slicing in 5G Heterogeneous Networks
The emergence of 5G networks, aimed at accommodating diverse service requirements such as enhanced Mobile Broadband (eMBB), Ultra-Reliable Low-Latency Communication (URLLC), and massive Machine-Type Communication (mMTC), has presented significant challenges in radio resource management and network slicing. In dynamic heterogeneous network systems, traditional heuristics and mathematical programming methods find it challenging to attain scalable multi-objective optimization while adhering to stringent service-specific restrictions. This research presents a unique two-stage Deep Reinforcement Learning (DRL) architecture that disaggregates the joint radio resource management problem into (i) slice admission and base-station assignment and (ii) resource block-level allocation within each gNodeB. Stage 1 utilizes an ensemble-based deep learning scheduler that chooses among many candidate slice-to-BS assignment strategies produced by diverse neural networks. While, Stage 2 employs Deep Q-Network (DQN) agents for the dynamic allocation of resource blocks to User Equipment (UE) at each gNodeB. The concept clearly integrates slice-specific Quality of Service (QoS) needs for eMBB, URLLC, and mMTC, embedding them into both the optimization constraints and the DRL reward framework. The suggested framework is executed in NS-3.35 utilizing the LTE-NR module with 3GPP TR 38.901 channel models and a heterogeneous macro/small-cell deployment. Comprehensive simulations under realistic traffic conditions demonstrate that, at elevated load, the proposed DRL framework diminishes average service delay by approximately 32%, enhances spectral efficiency by 28%, and reduces the dynamic energy consumption by approximately 24% with respect to the Greedy baseline at the same load, while preserving URLLC latency below 1 ms and URLLC reliability above 99.998% in our NS-3 scenarios, which is very close to—but does not fully reach—the strict 99.999% target.