Accès ouvert
2026
conference-paper
OpenAlex
Raphaela Baybas, Carlo D’Eramo, Philipp Brune
Adaptive Reward Design (ARD) is becoming a fundamental component for Reinforcement Learning (RL) agents, as they are deployed in increasingly complex settings where a single static reward across all phases of learning is rarely sufficient. Yet, ARD is rarely studied as a …
de
(code pays fourni par la source)
2026
article
OpenAlex
Jose-Luis Holgado-Alvarez, Gabriele Tiboni, Aryaman Reddi, Carlo D’Eramo
de
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Luca Ghisi, Jacopo Essenziale, Carlo D’Eramo, Matteo Luperto
Autonomous Racing has seen remarkable progress through deep Reinforcement Learning (RL), primarily for four-wheeled vehicles. However, motorbikes introduce substantially greater complexity due to the need to manage balance and lean angle, in addition to more reactive steering and throttle control, and a …
Accès ouvert
2026
preprint
OpenAlex
Luca Ghisi, Jacopo Essenziale, Carlo D’Eramo, Matteo Luperto
Autonomous Racing has seen remarkable progress through deep Reinforcement Learning (RL), primarily for four-wheeled vehicles. However, motorbikes introduce substantially greater complexity due to the need to manage balance and lean angle, in addition to more reactive steering and throttle control, and a …
it, de
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Zechu Li, Yufeng Jin, Xiaoyang Liu, Puze Liu et autres
Reinforcement learning (RL) has become a powerful paradigm for robot learning, particularly in sim-to-real settings, but its broader adoption remains limited by the engineering pipeline surrounding the algorithms. Building tasks, shaping rewards, and tuning hyperparameters require substantial expert effort, making RL workflows …
Accès ouvert
2026
preprint
OpenAlex
Zechu Li, Yufeng Jin, Xiaoyang Liu, Puze Liu et autres
Reinforcement learning (RL) has become a powerful paradigm for robot learning, particularly in sim-to-real settings, but its broader adoption remains limited by the engineering pipeline surrounding the algorithms. Building tasks, shaping rewards, and tuning hyperparameters require substantial expert effort, making RL workflows …
Accès ouvert
2026
preprint
OpenAlex
Noah Farr, Aryaman Reddi, Carlo D’Eramo, Jan Peters
Streaming reinforcement learning has emerged as an online learning paradigm that conforms to the restrictions of natural learning agents that process data incrementally, i.e. with a batch size of 1 and no replay buffer. While streaming RL has recently been shown to …
Accès ouvert
2026
preprint
OpenAlex
Noah Farr, Aryaman Reddi, Carlo D’Eramo, Jan Peters
Streaming reinforcement learning has emerged as an online learning paradigm that conforms to the restrictions of natural learning agents that process data incrementally, i.e. with a batch size of 1 and no replay buffer. While streaming RL has recently been shown to …
de
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Mahdi Kallel, Johannes Tölle, Ahmed Hendawy, Carlo D’Eramo
Standard supervised classification trains models to imitate the exact labels provided by a perfect oracle. This imitation happens in a single pass, restricting the model to a fixed compute budget even when inputs vary in complexity. Moreover, the rigid training objective forces …
Accès ouvert
2026
preprint
OpenAlex
Mahdi Kallel, Johannes Tölle, Ahmed Hendawy, Carlo D’Eramo
Standard supervised classification trains models to imitate the exact labels provided by a perfect oracle. This imitation happens in a single pass, restricting the model to a fixed compute budget even when inputs vary in complexity. Moreover, the rigid training objective forces …
de, ca
(code pays fourni par la source)
Accès ouvert
2026
article
OpenAlex
Arash Torabi Goodarzi, Waxenegger-Wilfing Günther, Carlo D’Eramo
Accès ouvert
2025
preprint
OpenAlex
Ahmed Hendawy, Henrik Metternich, Théo Vincent, Jan Peters et autres
The use of target networks is a popular approach for estimating value functions in deep Reinforcement Learning (RL). While effective, the target network remains a compromise solution that preserves stability at the cost of slowly moving targets, thus delaying learning. Conversely, using …