Trajectory-Oriented Policy Optimization with Sparse Rewards
Rattachement africain : cn. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
The mastery of deep reinforcement learning (DRL) proves challenging when the environmental reward signals are sparse. These limited rewards indicate whether the task has been partially or entirely accomplished, and the agent must conduct various exploration actions before the agent receives meaningful feedback. Therefore, most existing DRL exploration algorithms cannot acquire practical behavior policies within a reasonable term. This study introduces an RL method that leverages offline demonstration trajectories for a faster and more efficient online RL in environments with sparse rewards. The crucial insight we provide is to treat offline demonstration trajectories as guidance, rather than merely imitation, allowing our method to identify a policy with a distribution of state-action visitation that is marginally in line with offline demonstrations. A new trajectory distance based on maximum mean discrepancy (MMD) is presented and cast as a distance-constrained optimization problem. As a result, we demonstrate that the optimization problem can be streamlined by using a policy-gradient algorithm, incorporating rewards based on insights acquired from offline demonstrations. In this study, the proposed algorithm is evaluated across a navigation task with a discrete action space and two continuous locomotion control tasks. According to our experimental findings, our suggested algorithm offers significant advantages over baseline methods in terms of exploring diverse policy spaces and acquiring optimal policies.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Trajectory-Oriented Policy Optimization with Sparse Rewards
- Date Crossref
- 17/05/2024
- Éditeur
- IEEE
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
Beihang University pays non établi dans la noticeUniversité ou école supérieure
Beihang University.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.