יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics

הצגנו את TOPReward, שיטה לתוכנת פרגמנטים ספורדים ללמידת רובוטים. השיטה משתמשת במודלי וידאו-שפה מוכשרים כדי לספק חזרה צפויה ללא צורך בהכשרה.
תקציר מקורי באנגליתarXiv:2602.19313v2 Announce Type: replace-cross Abstract: General-purpose robot learning requires dense, instruction-conditioned feedback that can distinguish meaningful task progress from stalled, failed, or partially completed behavior. Yet obtaining such feedback at scale remains difficult, since existing approaches often rely on manual progress annotations, task-specific demonstrations, or reward models trained on curated robot datasets. We introduce TOPReward, a training-free progress reward method that probes pretrained Video-Language Models (VLMs) through their internal token probabilities rather than asking them to generate numerical progress values. Given a video prefix and a language instruction, TOPReward measures the model's likelihood that the instructed task has been complete
קרא במקור המקורי