יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

למידה מחדש ברמה התו: פיתוח תוצאות נאמנות תחת שינוי תפוצה

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift
אנו מציגים פרדיגמה של למידה מחדש ברמה התו, TOPL, שמטרתה לפתח תוצאות נאמנות תחת שינוי תפוצה. TOPL מצליח להשיג תוצאות טובות בתחום סיכום תיעודים ותרגום מכוני.
תקציר מקורי באנגליתarXiv:2607.17524v2 Announce Type: replace Abstract: We propose Token-Level Off-Policy Labeling (TOPL), an off-policy training paradigm that reframes post-training as a token-level correctness prediction task. Our key intuition is that by training the model to distinguish good and bad tokens in a response, we naturally guide the model towards generating good tokens, while avoiding the pitfalls that come with directly training the model to generate off-policy tokens. Experiments on document summarization tasks show that TOPL achieves strong out-of-distribution generalization across 11 datasets against a diverse set of sequence-level and token-level baselines. We further demonstrate that TOPL transfers effectively to machine translation, suggesting that its benefits generalize across differen
קרא במקור המקורי