כתבה
arXiv cs.AI ·
פורק: איפה המודל שונה את דעתו: חידושי גזרה ללמידת רפלקסיה מבנית
Fork Where the Model Changes Its Mind: Belief-Shift Branching for Tree-Structured Reinforcement Learning
מאמר חדש מציג חידוש בלמידת רפלקסיה מבנית, המשפר את הקרדיט לשלבי הלמידה. החידוש נקרא 'שינוי-דעה' ומאפשר למודלים לשנות את דעתם במהלך הלמידה.
תקציר מקורי באנגליתarXiv:2609.11061v1 Announce Type: new Abstract: Tree-structured rollouts give critic-free reinforcement learning with verifiable rewards (RLVR) step-level credit: fork a chain at an intermediate point, and sibling outcome differences estimate step value. Each fork adds sampling cost, so realistic budgets typically allow only a few forks per chain. A fork placed where the outcome is already largely settled yields siblings that mostly agree and provide almost no credit signal; hence, for a given tree size, where forks are placed largely determines how much step-level RL can gain. Most existing mainstream methods place forks by structure, such as fixed lengths, midpoints, and delimiters, or by next-token entropy. We formalize fork placement as locating the \emph{pivots} of the chain's value c
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית