כתבה
arXiv cs.LG ·
הסתברות-נורמליזציה של גבולות לאופטימיזציה של טעמים ישירים
Uncertainty-Normalized Margins for Direct Preference Optimization
מאמר זה עוסק בפיתוח שיטה לאופטימיזציה של טעמים ישירים, המשלבת גבולות של הסתברות-נורמליזציה. השיטה, הנקראת UNM-DPO, משתמשת במודלי Bradley-Terry עם קנה מידה משותף של רעש, ומציעה פתרון לבעיית הסתברות-נורמליזציה. המאמר כולל תיאור של השיטה, תיאור של המודלים שנעשה שימוש בהם, ותוצאות של ניסויים שנערכו כדי לבדוק את השיטה.
תקציר מקורי באנגליתarXiv:2609.38647v1 Announce Type: new Abstract: Direct preference optimization (DPO) models binary preferences through a Bradley-Terry model with a common noise scale, without explicitly accounting for preference strength or prompt-dependent uncertainty from human feedback. We introduce uncertainty-normalized margin DPO (UNM-DPO), which combines strength-dependent margins with a learned prompt scale. Motivated by a heteroskedastic Bradley-Terry model, we develop two training objectives. Both compare the implicit rewards of preferred and rejected responses, derived from response log-probability ratios to a reference policy. Advantage-only (AO) divides this reward difference by the prompt scale before subtracting the margin; whole-residual (WR) subtracts the margin before dividing by the sca
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית