Policy search via the signed derivative
Rattachement africain : us. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
We consider policy search for reinforcement learning: learning policy parameters, for some fixed policy class, that optimize performance of a system.In this paper, we propose a novel policy gradient method based on an approximation we call the Signed Derivative; the approximation is based on the intuition that it is often very easy to guess the direction in which control inputs affect future state variables, even if we do not have an accurate model of the system.The resulting algorithm is very simple, requires no model of the environment, and we show that it can outperform standard stochastic estimators of the gradient; indeed we show that Signed Derivative algorithm can in fact perform as well as the true (model-based) policy gradient, but without knowledge of the model.We evaluate the algorithm's performance on both a simulated task and two realworld tasks -driving an RC car along a specified trajectory, and jumping onto obstacles with an quadruped robot -and in all cases achieve good performance after very little training.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Policy search via the signed derivative
- Date Crossref
- 28/06/2009
- Éditeur
- Robotics: Science and Systems Foundation
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.