כתבה
arXiv cs.AI ·
לכיוון פעילות אופטימיזציה נפרדת: דפוסי השפעה של תענוג
Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner
אופטימיזציה של תענוג: דרך לדפוסי השפעה של תענוג שמפחיתה תגובות שנדחו ושומרת תגובות שנבחרו. ניתן לראות זאת ככלי לשפר את התוצאות של LLMs.
תקציר מקורי באנגליתarXiv:2604.18239v4 Announce Type: replace-cross Abstract: Preference optimization is widely used to align large language models (LLMs) with human preferences. However, many margin-based methods also suppress the chosen response when they try to suppress the rejected one, and there is no general way to prevent this across different objectives. We address this issue with a unified incentive-score decomposition of preference optimization, revealing that different objectives share the same local update directions and differ only in their scalar weights. This decomposition provides a common framework for analyzing objectives that were previously studied in separate settings. Building on this decomposition, by analyzing the dynamics of the chosen/rejected likelihoods, we identify the disentangle
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית