一种高斯过程的带参近似策略迭代算法

doi:10.3724/SP.J.1001.2013.04466

微信服务号

微信订阅号

2025年5月2日 0:09 星期五

首页 > 过刊浏览>2013年第24卷第11期 >2676-2686. DOI:10.3724/SP.J.1001.2013.04466

PDF HTML阅读 XML下载导出引用引用提醒

一种高斯过程的带参近似策略迭代算法
DOI:
                        10.3724/SP.J.1001.2013.04466
                    
CSTR:
                        
                    
作者:
                        傅启明傅启明
苏州大学 计算机科学与技术学院, 江苏 苏州 215006
在期刊界中查找
在百度中查找
在本站中查找
刘全刘全
苏州大学 计算机科学与技术学院, 江苏 苏州 215006;符号计算与知识工程教育部重点实验室吉林大学, 吉林 长春 130012
在期刊界中查找
在百度中查找
在本站中查找
伏玉琛伏玉琛
苏州大学 计算机科学与技术学院, 江苏 苏州 215006
在期刊界中查找
在百度中查找
在本站中查找
周谊成周谊成
苏州大学 计算机科学与技术学院, 江苏 苏州 215006
在期刊界中查找
在百度中查找
在本站中查找
于俊于俊
苏州大学 计算机科学与技术学院, 江苏 苏州 215006
在期刊界中查找
在百度中查找
在本站中查找

                    
作者单位:
作者简介:
通讯作者:
中图分类号:
基金项目:国家自然科学基金(61070223,61103045,61170020,61272005,61272244);江苏省自然科学基金(BK2012616);吉林大学符号计算与知识工程教育部重点实验室基金(93K172012K04)

Parametric Approximation Policy Iteration Algorithm Based on Gaussian Process

Author:

FU Qi-Ming
FU Qi-Ming
School of Computer Science and Technology, Soochow University, Suzhou 215006, China
在期刊界中查找
在百度中查找
在本站中查找
LIU Quan
LIU Quan
School of Computer Science and Technology, Soochow University, Suzhou 215006, China;Key Laboratory of Symbolic Computation and Knowledge Engineering Jilin University, Ministry of Education, Changchun 130012, China
在期刊界中查找
在百度中查找
在本站中查找
FU Yu-Chen
FU Yu-Chen
School of Computer Science and Technology, Soochow University, Suzhou 215006, China
在期刊界中查找
在百度中查找
在本站中查找
ZHOU Yi-Cheng
ZHOU Yi-Cheng
School of Computer Science and Technology, Soochow University, Suzhou 215006, China
在期刊界中查找
在百度中查找
在本站中查找
YU Jun
YU Jun
School of Computer Science and Technology, Soochow University, Suzhou 215006, China
在期刊界中查找
在百度中查找
在本站中查找

Affiliation:

Fund Project:

摘要

图/表

访问统计

参考文献

相似文献

引证文献

资源附件

文章评论

摘要:

在大规模状态空间或者连续状态空间中,将函数近似与强化学习相结合是当前机器学习领域的一个研究热点;同时,在学习过程中如何平衡探索和利用的问题更是强化学习领域的一个研究难点.针对大规模状态空间或者连续状态空间、确定环境问题中的探索和利用的平衡问题,提出了一种基于高斯过程的近似策略迭代算法.该算法利用高斯过程对带参值函数进行建模,结合生成模型,根据贝叶斯推理,求解值函数的后验分布.在学习过程中,根据值函数的概率分布,求解动作的信息价值增益,结合值函数的期望值,选择相应的动作.在一定程度上,该算法可以解决探索和利用的平衡问题,加快算法收敛.将该算法用于经典的Mountain Car 问题,实验结果表明,该算法收敛速度较快,收敛精度较好.

关键词:强化学习;策略迭代;高斯过程;贝叶斯推理;函数近似

Abstract:

In machine learning with large or continuous state space, it is a hot topic to combine the function approximation and reinforcement learning. The study also faces a very difficult problem of how to balance the exploration and exploitation in reinforcement learning. In allusion to the exploration and exploitation dilemma in the large or continuous state space, this paper presents a novel policy iteration algorithm based on Gaussian process in deterministic environment. The algorithm uses Gaussian process to model the action-value function, and in conjunction with generative model, obtains the posteriori distribution of the parameter vector of the action-value function by Bayesian inference. During the learning process, it computes the value of perfect information according to the posteriori distribution, and then selects the appropriate action with respect to the expected value of the action-value function. The algorithm achieves the balance between exploration and exploitation to certain extent, and therefore accelerates the convergence. The experimental results on the Mountain Car problem show that the algorithm has faster convergence rate and better convergence performance.

Key words:reinforcement learning;policy iteration;Gaussian process;Bayesian inference;function approximation

引用本文

傅启明,刘全,伏玉琛,周谊成,于俊.一种高斯过程的带参近似策略迭代算法.软件学报,2013,24(11):2676-2686

复制

文章指标

点击次数:
下载次数:
HTML阅读次数:
引用次数:

历史

收稿日期:2013-01-29
最后修改日期:2013-07-16
录用日期:
在线发布日期: 2013-11-01
出版日期:

微信服务号

微信订阅号

引用本文

分享

文章指标

历史

文章二维码

微信服务号

微信订阅号

引用本文

分享

微信扫一扫：分享

文章指标

历史

文章二维码