כתבה
arXiv cs.AI ·
מודלים סביבתיים: חשיבה מחדש על טיפול בפידבק בהסתכלות עצמית במסגרת חסכון
Environmental Feedback Modeling Matters: Rethinking Feedback Treatment in Agentic Hindsight Self-Distillation
במאמר זה, נחקר כיצד ניתן לשפר את יכולות האגנטים באמצעות חסכון סביבתי. המחברים מציגים פרקטיקה חדשה של SELF, שמשתמשת בפידבק סביבתי כדי לשפר את ההסתכלות העצמית של האגנט. התוצאות המוצגות במאמר זה מראות כי SELF משפרת את יכולות האגנטים באמצעות חסכון סביבתי.
תקציר מקורי באנגליתarXiv:2610.11384v1 Announce Type: new Abstract: Reinforcement learning is commonly used to train language agents in interactive environments, but cannot be directly applied when rewards are unavailable. Recent methods use environmental feedback as privileged context for hindsight self-distillation, but our analysis suggests that simply conditioning the teacher on feedback is insufficient, motivating us to rethink how environmental feedback is used in agentic self-distillation. Given that environmental feedback contains rich supervision for modeling how the environment responds to agent actions, we introduce \textit{agentic SElf-distilLation with environmental Feedback modeling} (SELF), a framework that jointly optimizes environmental feedback modeling and hindsight self-distillation. SELF
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית