Accès ouvert
2025
article
OpenAlex
Mahammad Humayoo, Gengzhong Zheng, Xiaoqing Dong, Wei Huang et autres
In numerous episodic reinforcement learning (RL) environments, SARSA-based methodologies are employed to enhance policies aimed at maximizing returns over long horizons. Traditional SARSA algorithms face challenges in achieving an optimal balance between bias and variation, primarily due to their dependence on a …
cn
(code pays fourni par la source)
Accès ouvert
2025
article
OpenAlex
Mahammad Humayoo, Gengzhong Zheng, Xiaoqing Dong, Liming Miao et autres
Off-policy learning exhibits greater instability when compared to on-policy learning in reinforcement learning (RL). The difference in probability distribution between the target policy (π) and the behavior policy (b) is a major cause of instability. High variance also originates from distributional mismatch. …
cn
(code pays fourni par la source)
Accès ouvert
2024
preprint
OpenAlex
Mahammad Humayoo
In numerous episodic reinforcement learning (RL) environments, SARSA-based methodologies are employed to enhance policies aimed at maximizing returns over long horizons. Traditional SARSA algorithms face challenges in achieving an optimal balance between bias and variation, primarily due to their dependence on a …
Accès ouvert
2024
preprint
OpenAlex
Mahammad Humayoo
Q-Learning is a fundamental off-policy reinforcement learning (RL) algorithm that has the objective of approximating action-value functions in order to learn optimal policies. Nonetheless, it has difficulties in reconciling bias with variance, particularly in the context of long-term rewards. This paper introduces …
Accès ouvert
2024
article
OpenAlex
Naveed Ur Rehman Junejo, Qingsheng Huang, Xiaoqing Dong, Chang Wang et autres
A variety of reasons have made it more difficult for educators and tutors to anticipate students’ performance. Numerous researchers have used various predictive models to identify students who may be at-risk of dropping out early. Additionally, these methods were used to forecast …
cn
(code pays fourni par la source)
2024
conference-paper
OpenAlex
Naveed Ur Rehman Junejo, Qingsheng Huang, Xiaoqing Dong, Chang Wang et autres
cn
(code pays fourni par la source)
Accès ouvert
2018
preprint
OpenAlex
Mahammad Humayoo
Off-policy learning exhibits greater instability when compared to on-policy learning in reinforcement learning (RL). The difference in probability distribution between the target policy ($π$) and the behavior policy (b) is a major cause of instability. High variance also originates from distributional mismatch. …
2018
conference-paper
OpenAlex
Mahammad Humayoo, Xueqi Cheng
Automatic selection of true explanatory variables and controlling fraction of false discovery rate (FDR) in the linear model has received considerable attention in machine learning. The ordered regularization is an important component of the linear model and plays a key role in …
cn
(code pays fourni par la source)
2014
conference-paper
OpenAlex
Mahammad Humayoo, Yanlong Zhai, Yan He, Bingqing Xu et autres
cn
(code pays fourni par la source)