כתבה
arXiv cs.LG ·
FAR: חידוש-מודע חזרה לפעולה לשיקום ושיפור רציף של פוליצי
FAR: Failure-Aware Retry for Test-Time Recovery and Continual Policy Improvement
FAR - פלטפורמה לחזרה לפעולה ושיפור רציף של פוליצי, המאפשרת לרובוטים ללמוד מכשלונים ולשפר את ביצועיהם. הפלטפורמה משתמשת בשיטת חידוש-מודע ובפעולות פרטורבציות קלות כדי לעודד התפשרות סביב נקודות הכשל. ניסויים בשני סביבות - סימולציה ועולם אמיתי - הראו ש-FAR משפרת את קצב ההצלחה ואת התוקפנות של הרובוטים, וכן משפרת את כושר הנתונים שלהם.
תקציר מקורי באנגליתarXiv:2607.01111v2 Announce Type: replace-cross Abstract: Robot policies inevitably encounter failures when deployed in real environments. Naive retries often repeat the same mistakes, while many existing recovery methods rely on human intervention. In this paper, we propose Failure-Aware Retry (FAR), a framework that enables robots to learn from previous failures at test time, adapt their behavior accordingly, and eventually complete the task autonomously. FAR combines Failure-Contrastive Preference Adaptation, which constructs preference learning data from failures to steer the policy away from previously unsuccessful behaviors, with lightweight action perturbations during retries to encourage local exploration. We further incorporate successful recovery trajectories into a training loop
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית