כתבה
arXiv cs.AI ·
בחירת זוגות מותאמת גרדיאנטים לאופטימיזציה אישית
Gradient-Aligned Pair Selection for Personalized Preference Optimization
GAP-DPO הוא אלגוריתם חדש לאופטימיזציה אישית של מודלי שפה גדולים. הוא משפר את איכות היצירה וההתאמה האישית על ידי בחירת זוגות מותאמת גרדיאנטים. האלגוריתם מוכיח את עצמו בניסויים על מודלים שונים, כולל LLaMA.
תקציר מקורי באנגליתarXiv:2610.00061v1 Announce Type: new Abstract: Personalizing large language models (LLMs) requires aligning generation behavior with user-specific preferences rather than aggregate quality. While Direct Preference Optimization (DPO) provides a stable framework for preference learning, its effectiveness in personalized settings critically depends on how preference pairs are selected. Existing approaches typically rely on heuristic criteria, such as likelihood-based extremes, which decouple optimization from explicit user utility and can lead to degraded personalization. We formalize personalized preference learning as a geometry-aligned optimization problem by analyzing the first-order interaction between gradients of expected user utility and DPO update directions. Our analysis reveals th
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית