כתבה
arXiv cs.AI ·
התבוננות עצמית: הפיכת חוויות אחר-העובדה לתפישה מוקדמת
Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
במאמר זה, המחברים מציגים שיטה חדשה לשיפור יכולות הצפייה של זיהוי סיביות, המשתמשת בחוויות אחר-העובדה. השיטה, הקרויה Self-Retrospection Distillation, מספקת תרגילים לזיהוי סיביות שמבוססים על חוויות אחר-העובדה, ומשפרות את יכולות הצפייה של הזיהוי. המחברים מציגים תוצאות של ניסויים שהראו שהשיטה החדשה משפרת את יכולות הצפייה של הזיהוי, ומציעה תרחיקי אף לפיתוח זיהוי סיביות טוב יותר.
תקציר מקורי באנגליתarXiv:2610.08077v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction. For group-relative objectives, however, this signal vanishes when all rollouts receive the same reward, even though their trajectories may reveal useful information about what the task requires and how the agent fails. We ask a complementary question: can hindsight teach an agent what it could have anticipated before acting? We introduce prospective learning, which uses post-hoc experience to supervise foresight predictions from the pre-interaction view, and instantiate it with Self-Retrospection Distillation (SRD). Intuitively, a completed trajectory reveals knowledge that would ha
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית