כתבה
arXiv cs.LG ·
Critic Architecture Matters: Dual vs. Unified Critics for Humanoid Loco-Manipulation
תקציר מקורי באנגליתarXiv:2606.11891v2 Announce Type: replace-cross Abstract: Multi-objective reinforcement learning for humanoid robots must coordinate locomotion and manipulation within a single policy. A natural design choice is whether to use a single (unified) critic that estimates the combined value of all objectives, or separate (dual) critics with disjoint reward signals. We compare the two on the Unitree G1 humanoid (23 active DoF, of which 17 are policy-controlled) in NVIDIA Isaac Lab, training loco-manipulation policies through sequential curricula that progress from stationary reaching to walking with variable-orientation targets. Under a matched compute budget, the dual-critic run reaches targets 3.5x faster (6.5 vs. 22.6 simulation steps), achieves 2x higher throughput (14.3 vs. 7.0 validated re
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית