כתבה
arXiv cs.AI ·
למידה מעדויות רעשיות: גישה חצי-סופרוויזד לאופטימיזציה של עדות ישירה
Learning from Noisy Preferences: A Semi-Supervised Learning Approach to Direct Preference Optimization
מאמר חדש מציע גישה חצי-סופרוויזד לאופטימיזציה של עדות ישירה, כדי ללמוד מעדויות רעשיות. השיטה משתמשת בקבוצות נקיות ולא-נקיות כדי לשפר את הביצועים.
תקציר מקורי באנגליתarXiv:2604.24952v3 Announce Type: replace-cross Abstract: Human visual preferences are inherently multi-dimensional, encompassing aesthetics, detail fidelity, and semantic alignment. However, existing datasets provide only single, holistic annotations, resulting in severe label noise: images that excel in some dimensions but are deficient in others are simply marked as winner or loser. We theoretically demonstrate that compressing multi-dimensional preferences into binary labels generates conflicting gradient signals that misguide Diffusion Direct Preference Optimization (DPO). To address this, we propose Semi-DPO, a semi-supervised approach that treats consistent pairs as clean labeled data and conflicting ones as noisy unlabeled data. Our method starts by training on a consensus-filtered
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית