Parametric Approximation Policy Iteration Algorithm Based on Gaussian Process

doi:10.3724/SP.J.1001.2013.04466

微信服务号

微信订阅号

2025-5-11- 8

Home > Archive>Volume 24, Issue 11, 2013 >2676-2686. DOI:10.3724/SP.J.1001.2013.04466

PDF HTML XML Export Cite reminder

Parametric Approximation Policy Iteration Algorithm Based on Gaussian Process
DOI:
                        10.3724/SP.J.1001.2013.04466
                    
Author:
                        FU Qi-MingFU Qi-Ming
School of Computer Science and Technology, Soochow University, Suzhou 215006, China
Find this author on CNKI
Find this author on BaiDu
Search for this author on this site
LIU QuanLIU Quan
School of Computer Science and Technology, Soochow University, Suzhou 215006, China;Key Laboratory of Symbolic Computation and Knowledge Engineering Jilin University, Ministry of Education, Changchun 130012, China
Find this author on CNKI
Find this author on BaiDu
Search for this author on this site
FU Yu-ChenFU Yu-Chen
School of Computer Science and Technology, Soochow University, Suzhou 215006, China
Find this author on CNKI
Find this author on BaiDu
Search for this author on this site
ZHOU Yi-ChengZHOU Yi-Cheng
School of Computer Science and Technology, Soochow University, Suzhou 215006, China
Find this author on CNKI
Find this author on BaiDu
Search for this author on this site
YU JunYU Jun
School of Computer Science and Technology, Soochow University, Suzhou 215006, China
Find this author on CNKI
Find this author on BaiDu
Search for this author on this site

                    
Affiliation:
Clc Number:
Fund Project:

Article

Figures

Metrics

Reference

Cited by

Materials

Comments

Abstract:

In machine learning with large or continuous state space, it is a hot topic to combine the function approximation and reinforcement learning. The study also faces a very difficult problem of how to balance the exploration and exploitation in reinforcement learning. In allusion to the exploration and exploitation dilemma in the large or continuous state space, this paper presents a novel policy iteration algorithm based on Gaussian process in deterministic environment. The algorithm uses Gaussian process to model the action-value function, and in conjunction with generative model, obtains the posteriori distribution of the parameter vector of the action-value function by Bayesian inference. During the learning process, it computes the value of perfect information according to the posteriori distribution, and then selects the appropriate action with respect to the expected value of the action-value function. The algorithm achieves the balance between exploration and exploitation to certain extent, and therefore accelerates the convergence. The experimental results on the Mountain Car problem show that the algorithm has faster convergence rate and better convergence performance.

Key words:reinforcement learning;policy iteration;Gaussian process;Bayesian inference;function approximation

Get Citation

傅启明,刘全,伏玉琛,周谊成,于俊.一种高斯过程的带参近似策略迭代算法.软件学报,2013,24(11):2676-2686

Copy

Article Metrics

Abstract:
PDF:
HTML:
Cited by:

History

Received:January 29,2013
Revised:July 16,2013
Adopted:
Online: November 01,2013
Published:

You are the first2043738Visitors
Copyright: Institute of Software, Chinese Academy of Sciences Beijing ICP No. 05046678-4
Address：4# South Fourth Street, Zhong Guan Cun, Beijing 100190,Postal Code：100190
Phone：010-62562563 Fax：010-62562533 Email：jos@iscas.ac.cn
Technical Support：Beijing Qinyun Technology Development Co., Ltd.

Beijing Public Network Security No. 11040202500063

微信服务号

微信订阅号

Get Citation

Share

微信扫一扫：分享

Article Metrics

History