כתבה
arXiv cs.AI ·
הבחנה עתידית: הכשרה עצמית של למידת ריפוד על ידי פערי תחזית-אמת
Prospective Hindsight: Self-Calibrating Reinforcement Learning via Prediction-Reality Gaps
אופציונלי של למידת ריפוד שמסתגל לעצמו על ידי פערי תחזית-אמת. המאמר עוסק בפיתוח של טכניקה חדשה שמספקת סיגנל נוסף ללמידת ריפוד, המבוסס על הפער בין התחזית העתידית של האגן (לפני קבלת הפידבק) לבין הבדיקה הרטרוספקטיבית (אחרי קבלת הפידבק).
תקציר מקורי באנגליתarXiv:2610.02740v1 Announce Type: cross Abstract: Reinforcement learning for long-horizon agents relies on purely retrospective training signals: credit is assigned only after observing environmental consequences, leaving the agent's belief at action time invisible to the gradient. We introduce Prospective Hindsight (PH), a self-calibrating training principle that augments any retrospective base method with a signal derived from the gap between the agent's prospective prediction (before feedback) and the retrospective evaluation (after feedback). This per-rollout surprise identifies samples where the agent's self-model is most inaccurate and amplifies their gradient contribution through a stop-gradient surprise-weighted advantage. Since the prospective predictor shares parameters with the
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית