יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

פירוק יחסי-מדידה-לא-אופציונלי: תחומי אמון-אנטרופי-מורכבים ללמידת-ראשונה-אסינכרונית

Deconstructing Off-Policy Ratios: Entropy-Normalized Trust Regions for Asynchronous Reinforcement Learning
אנשי מדע פיתחו טכנולוגיה חדשה ללמידת-ראשונה-אסינכרונית, המשפרת את יעילות הלמידה ומונעת קריסת-מדל.
תקציר מקורי באנגליתarXiv:2607.22186v5 Announce Type: replace Abstract: Asynchronous reinforcement learning (RL) accelerates large language model (LLM) post-training by overlapping rollout generation with policy optimization, but the resulting stale, off-policy data destabilizes optimization and can cause policy collapse. Existing methods gate tokens by ratio magnitude alone, applying one threshold at every position. We show that the ratio's natural scale is set by token entropy, so deviations from mid-trajectory weight updates stay within this scale and carry genuine exploration. We further identify an overlooked low-entropy regime that breaks this scaling, where a near-zero probability amplifies train--inference mismatch into noise far beyond what the local entropy admits. A magnitude threshold admits this
קרא במקור המקורי