An Adaptive Exploration-Oriented Multi-Agent Co-Evolutionary Method Based on MATD3
Rattachement africain : cn. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
As artificial intelligence continues to evolve, reinforcement learning (RL) has shown remarkable potential for solving complex sequential decision problems and is now applied in diverse areas, including robotics, autonomous vehicles, and financial analytics. Among the various RL paradigms, multi-agent reinforcement learning (MARL) stands out for its ability to manage cooperative and competitive interactions within multi-entity systems. However, mainstream MARL algorithms still face critical challenges in training stability and policy generalization due to factors such as environmental non-stationarity, policy coupling, and inefficient sample utilization. To mitigate these limitations, this study introduces an enhanced algorithm named MATD3_AHD, developed by extending the MATD3 framework, which integrates TD3 and MADDPG principles. The goal is to improve the learning efficiency and overall policy effectiveness of agents operating in complex environments. The proposed method incorporates three key mechanisms: (1) an Adaptive Exploration Policy (AEP), which dynamically adjusts the perturbation magnitude based on TD error to improve both exploration capability and training stability; (2) a Hierarchical Sampling Policy (HSP), which enhances experience utilization through sample clustering and prioritized replay; and (3) a Dynamic Delayed Update (DDU), which adaptively modulates the actor update frequency based on critic network errors, thereby accelerating convergence and improving policy stability. Experiments conducted on multiple benchmark tasks within the Multi-Agent Particle Environment (MPE) demonstrate the superior performance of MATD3_AHD compared to baseline methods such as MADDPG and MATD3. The proposed MATD3_AHD algorithm outperforms baseline methods—by an average of 5% over MATD3 and 20% over MADDPG—achieving faster convergence, higher rewards, and more stable policy learning, thereby confirming its robustness and generalization capability.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- An Adaptive Exploration-Oriented Multi-Agent Co-Evolutionary Method Based on MATD3
- Date Crossref
- 26/10/2025
- Éditeur
- MDPI AG
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
China University of Mining and Technology pays non établi dans la noticeUniversité ou école supérieure
-
Hua Yuan Group (China) pays non établi dans la noticeEntreprise
-
China Academy of Railway Sciences pays non établi dans la noticeStructure de recherche
-
Institute of Intelligent Mining and Robotics pays non établi dans la noticeStructure de recherche
-
School of Mechanical and Electrical Engineering pays non établi dans la noticeUniversité ou école supérieure
-
Ltd. Beijing Huatie Information Technology Co. pays non établi dans la noticeEntreprise
-
Signal & Communication Research Institute pays non établi dans la noticeStructure de recherche
China University of Mining and Technology, Hua Yuan Group (China) et China Academy of Railway Sciences, avec 4 autres affiliations.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.