כתבה
arXiv cs.AI ·
שביעות רצון אקסיומטית של תגמולים ליניאריים
Axiom Satisfiability of Linear Rewards in Alignment
חוקרים פיתחו שיטה חדשה להתאמת תגמולים ליניאריים עם ערכים אנושיים. השיטה מאפשרת לקבוע תגמולים באופן שיתאים לערכים האנושיים, תוך שמירה על יעילות ודיוק. המחקר מציג תוצאות חיוביות בניסויים על נתונים מלאכותיים ואמיתיים.
תקציר מקורי באנגליתarXiv:2610.06892v1 Announce Type: cross Abstract: Learning from human preference data is the dominant route to aligning language models with human values. In linear social choice, where rewards are linear in a fixed feature representation of prompt-response pairs, Ge et al.[2024] show that fitting such a reward by minimizing any non-decreasing convex loss, including BTL, fails PO and PMC. Moreover, no rule that reads only the majority relation can satisfy PO once the output is required to be linearly induced. We ask what it costs to enforce these axioms anyway. To this end, we relax the linear model to allow per-candidate slack. We compute the relaxed linear reward with the smallest total slack that satisfies the axioms with a margin $\eta$, the minimum required difference between two rewa
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית