יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

למידה מעדויות רעשיות: גישה חצי-סופרוויזד לאופטימיזציה של עדות ישירה

Learning from Noisy Preferences: A Semi-Supervised Learning Approach to Direct Preference Optimization
מאמר חדש מציע גישה חצי-סופרוויזד לאופטימיזציה של עדות ישירה, כדי ללמוד מעדויות רעשיות. השיטה משתמשת בקבוצות נקיות ולא-נקיות כדי לשפר את הביצועים.
תקציר מקורי באנגליתarXiv:2604.24952v3 Announce Type: replace-cross Abstract: Human visual preferences are inherently multi-dimensional, encompassing aesthetics, detail fidelity, and semantic alignment. However, existing datasets provide only single, holistic annotations, resulting in severe label noise: images that excel in some dimensions but are deficient in others are simply marked as winner or loser. We theoretically demonstrate that compressing multi-dimensional preferences into binary labels generates conflicting gradient signals that misguide Diffusion Direct Preference Optimization (DPO). To address this, we propose Semi-DPO, a semi-supervised approach that treats consistent pairs as clean labeled data and conflicting ones as noisy unlabeled data. Our method starts by training on a consensus-filtered
קרא במקור המקורי