Convergence rate comparison of two data-driven algorithms to stochastic LQR problems
Rattachement africain : cn. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
This paper investigates the linear quadratic regulation (LQR) problem for discrete-time stochastic systems (DTSSs) with state-dependent multiplicative noise. Policy iteration (PI) and value iteration (VI) algorithms are proposed to solve the generalised algebraic Riccati equation (GARE) corresponding to the stochastic LQR (SLQR) problem. Furthermore, based on the proposed PI and VI algorithms, a comparative analysis of their convergence rates is conducted. Additionally, when the system dynamics are completely unknown, this paper introduces online model-free versions of the PI and VI algorithms using reinforcement learning (RL) technology to find the optimal control strategy for the SLQR problem. Finally, numerical simulations validate the feasibility of the proposed algorithms and the correctness of the theoretical results.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Convergence rate comparison of two data-driven algorithms to stochastic LQR problems
- Date Crossref
- 14/11/2025
- Éditeur
- Informa UK Limited
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.