כתבה
arXiv cs.AI ·
Large Reward Models: Generalizable Online Robot Reward Generation with Vision-Language Models
תקציר מקורי באנגליתarXiv:2603.16065v3 Announce Type: replace-cross Abstract: Reinforcement Learning (RL) has shown strong potential for improving robotic manipulation policies, yet its practical use remains bottlenecked by the difficulty of specifying reward functions that are both semantically meaningful and reusable across tasks. In this paper, we propose Large Reward Models (LRMs), a framework that adapts foundation VLMs into frame-level reward generators for robot policy refinement. We specialize a state-of-the-art VLM on a multi-source dataset spanning real-world robot trajectories, human-object interactions, and simulated manipulation environments. Unlike prior approaches that mainly evaluate trajectories post-hoc, LRMs expose multiple reward interfaces from visual observations: progress estimation, ta
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית