Aller au contenu principal
Profil bibliographique

Mahammad Humayoo

Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.

9Publications signalées
32Citations signalées
6Affiliations récentes

Les institutions déclarées

Les domaines associés

Reinforcement Learning in RoboticsOnline Learning and AnalyticsCOVID-19 diagnosis using AIAdvanced Bandit Algorithms ResearchStatistical Methods and Inference

Les publications récentes

Accès ouvert 2025 article OpenAlex

Segmenting Action-Value Functions over Time Scales in SARSA via TD(Δ)

Mahammad Humayoo, Gengzhong Zheng, Xiaoqing Dong, Wei Huang et autres

In numerous episodic reinforcement learning (RL) environments, SARSA-based methodologies are employed to enhance policies aimed at maximizing returns over long horizons. Traditional SARSA algorithms face challenges in achieving an optimal balance between bias and variation, primarily due to their dependence on a …

cn (code pays fourni par la source)

1 citation Algorithms
Accès ouvert 2025 article OpenAlex

Relative importance sampling for off-policy actor-critic in deep reinforcement learning

Mahammad Humayoo, Gengzhong Zheng, Xiaoqing Dong, Liming Miao et autres

Off-policy learning exhibits greater instability when compared to on-policy learning in reinforcement learning (RL). The difference in probability distribution between the target policy (π) and the behavior policy (b) is a major cause of instability. High variance also originates from distributional mismatch. …

cn (code pays fourni par la source)

3 citations Scientific Reports
Accès ouvert 2024 preprint OpenAlex

Segmenting Action-Value Functions Over Time-Scales in SARSA via TD($Δ$)

Mahammad Humayoo

In numerous episodic reinforcement learning (RL) environments, SARSA-based methodologies are employed to enhance policies aimed at maximizing returns over long horizons. Traditional SARSA algorithms face challenges in achieving an optimal balance between bias and variation, primarily due to their dependence on a …

0 citations arXiv (Cornell University)
Accès ouvert 2024 preprint OpenAlex

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition

Mahammad Humayoo

Q-Learning is a fundamental off-policy reinforcement learning (RL) algorithm that has the objective of approximating action-value functions in order to learn optimal policies. Nonetheless, it has difficulties in reconciling bias with variance, particularly in the context of long-term rewards. This paper introduces …

0 citations arXiv (Cornell University)
Accès ouvert 2024 article OpenAlex

SAPPNet: students’ academic performance prediction during COVID-19 using neural network

Naveed Ur Rehman Junejo, Qingsheng Huang, Xiaoqing Dong, Chang Wang et autres

A variety of reasons have made it more difficult for educators and tutors to anticipate students’ performance. Numerous researchers have used various predictive models to identify students who may be at-risk of dropping out early. Additionally, these methods were used to forecast …

cn (code pays fourni par la source)

19 citations Scientific Reports

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.