When Adding Arms Breaks Legs: Exploration Noise, Not Shared-Trunk Gradients, Governs the Collapse of From-Scratch Whole-Body RL for Humanoid Locomotion
Rattachement africain : jp. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
We study whether a single reinforcement-learning (RL) policy controlling all joints of ahumanoid (whole-body RL) is a viable alternative to the prevailing modular design (apretrained lower-body locomotion policy with a separately controlled upper body) forcommanded velocity-height tracking on a Unitree G1. Working entirely on a single consumerGPU (RTX 4060 Ti, 16 GB) with Isaac Lab and the AGILE framework, we report a sharp,reproducible failure mode: extending the action space from the 12 leg joints to all 29 bodyjoints causes from-scratch training to collapse -- episodes terminate within ~2 steps andnever recover -- even when the arms are physically frozen. Through controlled ablations werule out code defects, action scaling, data-augmentation symmetry, observation contamination,and the harness curriculum. The natural hypothesis -- gradient interference in the shared MLPtrunk -- we test directly and reject: decomposing the per-update trunk gradient into leg-headand upper-head contributions shows them indistinguishable between the collapsing and learningregimes, and an actor with physically separate trunks collapses identically. A clean 2x2(trunk topology x noise) isolates the controlling variable to the policy's initialexploration noise sigma_0; reducing it restores stable learning in both architectures, with asharp, seed-dependent dose-response. A compute-matched pilot finds the modular design reachesa deployable policy under an equal budget while from-scratch whole-body RL does not (success1.0 vs. 0.25), reported cautiously with explicit confounds. Our contribution is areproducible, single-GPU diagnosis that directly measures and falsifies the most citedcandidate mechanism, attributing the failure to exploration scale instead, plus a releasedevaluation harness and a confound-controlled protocol. Code & data:- Harness, configs, logs, paper source: https://gitlab.com/KyulingLee/sim_rl_data_factory_nvidia- Diagnostic probes (gradient decomposition, separate-trunk actor, ratio/LR probes), WBC-AGILE fork branch agile-h1-phase1: https://gitlab.com/KyulingLee/wbc-agile-h1
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.