כתבה
arXiv cs.LG ·
PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots
תקציר מקורי באנגליתarXiv:2610.01260v1 Announce Type: cross Abstract: Quadrupedal locomotion requires balancing conflicting objectives such as command tracking, stability, and energy efficiency, yet conventional reinforcement learning (RL) hardcodes these priorities into a fixed scalar reward at training time. We present PROMO (Preference-Conditioned Multi-Objective Reinforcement Learning), a semantic multi-objective approach that makes this trade-off an explicit runtime input to a single locomotion policy. PROMO conditions the policy on deployment facing preferences while keeping embodiment-specific locomotion priors fixed, thereby separating operator intent from reward shaping terms required for viable gait generation. Compared with fixed-objective controllers, multi-objective baselines, and independently t
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית