יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

חשיבה מחדש על שיפוט-בהסתברות בלמידת-רב-מודלים

Rethinking Probability-Based Reinforcement Learning From Posterior Concentration
במאמר זה, המחברים חוקרים את האפקט של קונצנטרציה פוסטריירית בלמידת-רב-מודלים. הם מציגים פרקטיקה חדשה ללמידת-רב-מודלים, שמטרתה לשפר את יעילות הטקסט ואת יציבות האופטימיזציה. הם גם מציגים תוצאות מעבדה, שמציגות את יעילות הפרקטיקה החדשה.
תקציר מקורי באנגליתarXiv:2610.01458v2 Announce Type: replace Abstract: Verifier-free reinforcement learning with probability-based rewards offers a promising way to train LLMs on general reasoning tasks where external verifiers are unavailable. Yet the reliability of these rewards, especially in long-horizon reasoning, remains underexplored. This work identifies a length-dependent failure mode of probability rewards, which we call the Posterior Concentration Phenomenon (PCP). We show that the probability of a reference answer conditioned on a reasoning trace often collapses to a low-variance interval as the trace becomes lengthy. This phenomenon results in nearly indistinguishable rewards, which, under GRPO-based settings, makes probability-based policy optimization unstable and inefficient. Motivated by thi
קרא במקור המקורי