כתבה
arXiv cs.AI ·
לימוד חיזוק עם הסתברות בלי וודיקטור
Rethinking Probability-Based Reinforcement Learning From Posterior Concentration
חוקרים זיהו תופעה של ריכוז אחורי בלימוד חיזוק עם הסתברות. הם הציעו מסגרת חדשה, RLCPR, שמשפרת יציבות ויעילות. RLCPR עובדת טוב יותר מבסיסי RL הקיימים.
תקציר מקורי באנגליתarXiv:2610.01458v1 Announce Type: new Abstract: Verifier-free reinforcement learning with probability-based rewards offers a promising way to train LLMs on general reasoning tasks where external verifiers are unavailable. Yet the reliability of these rewards, especially in long-horizon reasoning, remains underexplored. This work identifies a length-dependent failure mode of probability rewards, which we call the Posterior Concentration Phenomenon (PCP). We show that the probability of a reference answer conditioned on a reasoning trace often collapses to a low-variance interval as the trace becomes lengthy. This phenomenon results in nearly indistinguishable rewards, which, under GRPO-based settings, makes probability-based policy optimization unstable and inefficient. Motivated by this, w
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית