כתבה
arXiv cs.CL ·
GAW-PO: אופטימיזציה של תחביר על פי גרדיאנטים מאולץ
GAW-PO: Preference Optimization with Gradient-Aligned Token Weights
GAW-PO מאפשר אופטימיזציה של תחביר על פי גרדיאנטים מאולץ, כדי לשפר את תפקוד ה-LLM. השיטה משתמשת בגרדיאנטים כדי להעניק עדיפות לטקסט המועדף על פני טקסט שנדחה.
תקציר מקורי באנגליתarXiv:2610.01511v2 Announce Type: replace Abstract: Most preference optimization methods, such as Direct Preference Optimization (DPO), apply preference supervision at the response level, although autoregressive language models are optimized token by token. As a result, all tokens in a rejected response contribute to the negative training signal, including tokens that may encode behavior that is useful for the preferred response. We introduce GAW-PO, a gradient-aligned token reweighting method for DPO that estimates, for each rejected token, whether penalizing it would interfere with the preferred update directions. Tokens whose gradients are strongly aligned with the preferred behavior receive a weaker negative contribution, while conflicting tokens retain a stronger penalty. Our method a
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית