יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

התקדמות בהערכת קרדיט להסתברות ב-RLVR דרך אופטימיזציה של פרוקסימלי הסתברות

Advancing Entropy-Level Credit Assignment in RLVR via Proximal Entropy Policy Optimization
אנו מציגים טכניקה חדשה להערכת קרדיט להסתברות ב-RLVR, המשתמשת באופטימיזציה של פרוקסימלי הסתברות. הטכניקה שלנו, PEPO, משתמשת במדד חדש של הסתברות פרוקסימלי, המתאים להסתברות הסמוכה של כל תווית. PEPO משפרת את הביצועים של GRPO ובסיסים אחרים במתמטיקה ובהסתברות.
תקציר מקורי באנגליתarXiv:2609.39402v1 Announce Type: new Abstract: Value-model-free RLVR methods such as GRPO assign uniform advantages to all tokens in a rollout, ignoring that tokens contribute unequally. Recent methods use token entropy as an importance proxy but compute it globally across the batch, conflating importance with prompt difficulty and positional trends. We argue that importance should instead be measured relative to the local context of each token. We introduce proximal entropy, a local measure of token importance relative to neighboring tokens, and prove it is invariant to both confounders. Proximal Entropy Policy Optimization (PEPO) uses it to weight per-token advantages and outperforms GRPO and entropy-based baselines on mathematical reasoning across Qwen3-1.7B, Qwen3-4B, and Llama-3.2-3B
קרא במקור המקורי