Abstract:The issue of QoS (quality of service) provisioning for adaptive multimedia in wireless communication networks is considered. A reinforcement learning based online adaptive bandwidth allocation optimization algorithm is proposed. First, an event-driven stochastic switching model is introduced to formulate the adaptive bandwidth allocation problem as a constrained continuous-time Markov decision problem. Then, an online optimization algorithm that combines policy gradient estimation by learning and stochastic approximation is derived. This algorithm can online handle the constrained optimization problem efficiently without explicit knowledge of the underlying system parameters. Moreover, this algorithm does not require the computation of performance potentials or other related quantities (e.g. Q-factors), which is necessary in previous schemes, and therefore saves computational cost significantly. Simulation results demonstrate the effectiveness of the proposed algorithm.