כתבה
arXiv cs.LG ·
MInTRL: שיפור למידת חיזוק
MInTRL: Off-policy Intervention can boost On-policy RL
MInTRL היא שיטה חדשה לשיפור למידת חיזוק. היא מאפשרת הרחבת החיפוש מעבר לנתוני האימון הנוכחיים, תוך שמירה על הטבעיות של המדיניות. MInTRL משתמשת בהתערבויות מקומיות ונדירות כדי לשפר את הכיסוי.
תקציר מקורי באנגליתarXiv:2609.12419v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards is typically performed on-policy, keeping training data close to the current policy but limiting learning to trajectories that the policy can discover itself. Off-policy methods such as supervised fine-tuning, on the other hand, can leverage external knowledge beyond the base model's capabilities, but may suffer from large distribution shift. The key challenge is thus to expand exploration without sacrificing learnability. In this work, we introduce Minimal Intervention Reinforcement Learning (MInTRL), which expands the exploration frontier through sparse, local interventions in otherwise on-policy rollouts. During generation, a judge-intervention policy periodically reviews the current policy's
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית