יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

העשרה בלמידת תגמול

Understanding Enrichment in Reinforcement Learning
חוקרים בודקים את השפעת ההעשרה על למידת תגמול. הם מציגים מנגנון חדש לתיקון משקלים ומיישמים אותו על מודל Qwen3-1.7B.
תקציר מקורי באנגליתarXiv:2610.02846v1 Announce Type: new Abstract: When rewards are sparse, reinforcement learning with verifiable rewards (RLVR) often uses hints or intermediate guidance to generate more successful rollouts. This enrichment biases policy-gradient updates unless corrected via importance weights, but existing methods omit correction or truncate importance weights in order to avoid the high variance of correction. It thus remains unclear what exactly is gained or lost in RLVR by correcting enriched rollouts. We show, mathematically, that omitted or truncated correction implicitly reweights the defined reward, and we decompose the resulting gradient error into scale, rotation, and variance to explain their distinct effects on learning. To make correction practical, we develop a novel sequential
קרא במקור המקורי