יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

COPC: תיקון רפלקסיבי קשור לקורקטיבי רפלקסיבי ללמידת רפלקסיבי של LLM

COPC: Coupled Off-Policy Correction for Asynchronous LLM Reinforcement Learning
COPC: תיקון רפלקסיבי קשור לקורקטיבי רפלקסיבי ללמידת רפלקסיבי של LLM. תיקון זה נועד לתקן טרקטוריות ישנות של LLM. COPC משתמש בשיטה של תיקון רפלקסיבי קשור, שבה נעשה שימוש בשיטת תיקון רפלקסיבי קשור כדי לתקן טרקטוריות ישנות של LLM.
תקציר מקורי באנגליתarXiv:2610.09597v1 Announce Type: new Abstract: Asynchronous RL accelerates large language model post-training by decoupling rollout generation from optimization, but trains on stale trajectories. Existing methods primarily correct token-level policy mismatch through importance-ratio control in the actor objective. We show that this \emph{policy-side correction} alone is insufficient: advantage estimates also inherit mismatch from behavior-policy continuations, which we term \emph{advantage staleness}. We derive exact bias and variance decompositions for a general two-channel actor update, revealing nonseparable coupling between policy-weight and advantage-estimation errors: their interaction induces multiplicative bias terms, while squared policy weights amplify advantage uncertainty in g
קרא במקור המקורי