כתבה
arXiv cs.LG ·
A Balanced Data Diet: Addressing the Exploration Bottleneck in Mega-Scale RL for Robot Control
תקציר מקורי באנגליתarXiv:2610.12465v1 Announce Type: cross Abstract: General-purpose robots must perform a wide range of tasks from agile locomotion to dexterous manipulation. While sim-to-real reinforcement learning (RL) has proven to be a useful tool for this goal, current RL pipelines depend on engineering-heavy, per-task structural priors such as shaped rewards and demonstrations. Recent work has shown that diverse simulator resets, combined with massively parallel simulation, can alleviate much of this engineering burden on several manipulation problems. However, we find that naively scaling this paradigm to more precise or dynamic problems remains non-trivial. While simulator resets can help with exploration, uniformly sampling over this distribution wastes a growing fraction of learning experience on
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית