יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

למידה מהעתיד הקרוב: תפיסה זמנית של רווחי עצמי

Learning from the Near Future: Temporal Self-Distillation for RLVR
אופציוני רווחי עצמי עם תפיסה זמנית של רווחי עצמי משפרים את אופציוני המדידה וההפרדה של המדלים.
תקציר מקורי באנגליתarXiv:2604.20733v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) is a core post-training recipe for reasoning models, yet pure on-policy learning can be inefficient when useful trajectories are difficult to discover or exploration narrows. Existing self-guided approaches largely reuse capability already available to the current or earlier learner. We instead ask whether learning can also make use of capabilities that emerge later in training: can a model learn from its own future self? We introduce temporal self-distillation, in which a policy receives guidance from a stronger later checkpoint of itself. We hypothesize that the most useful temporal teacher need not be the strongest one: a teacher must provide sufficiently new capability while remain
קרא במקור המקורי