יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

KL Regularization באופטימיזציה של מדיניות קבוצה

When KL Regularization Misfires in Group Policy Optimization
KL Regularization יכולה לכשל באופטימיזציה של מדיניות קבוצה. מחקר זה בוחן שבעה מצבי כישלון אפשריים ביחסים בין KL ותגמולים. הוצעה Zero-Sum Calibrated Policy Optimization (ZCPO) כפתרון.
תקציר מקורי באנגליתarXiv:2610.12161v1 Announce Type: new Abstract: Why does removing reference-policy KL regularization sometimes improve group policy optimization? This motivates studying how reference-policy information should enter group-relative updates. We analyze seven potential failure modes in the interactions between KL and rewards: residual KL updates after reward clipping, after gradient cancellation, and in groups with identical rewards; KL growth with response length and an imbalance in its relative contribution; KL concentration on a small number of tokens; and sampling noise when k1 is incorporated into rewards. We propose Zero-Sum Calibrated Policy Optimization (ZCPO), which uses relative drift measured by conditional KL to calibrate within-group reward coefficients and integrates them into t
קרא במקור המקורי