כתבה
arXiv cs.LG ·
לגשר על הפער בין תקיפה אופלינית והתאמה איטרטיבית על ידי תרגום טעמים
Towards Bridging the Gap Between Offline and Iterative Alignment via Preference Distillation
מאמר חדש עוסק בפיתוח שיטה להתאמה של מודלי שפה, על ידי תרגום טעמים. השיטה, DP3O, משתמשת במודלים עזר כדי ללמוד טעמים ואז להפיץ את הידע לאופטימיזציה של פוליצי. המאמר כולל ניסויים שונים והוכחות תאורטיות.
תקציר מקורי באנגליתarXiv:2609.06893v1 Announce Type: cross Abstract: Direct preference optimization DPO is a promising offline approach for aligning large language models (LLMs) due to its simplicity, computational efficiency, and implicit modeling of human preferences. Interestingly, iterative extensions of DPO have achieved stronger performance on academic benchmarks, raising two key questions: (i) Why do iterative methods generally outperform offline ones? (ii) Can their advantages be incorporated into offline alignment? To answer the first question, our controlled experiments reveal that the explicit preference model, additionally introduced in the iterative procedure, is a key factor behind its superiority over offline methods. This insight leads us to answer the second question affirmatively and propos
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית