כתבה
arXiv cs.LG ·
משקלות למילים: ביטוי ועריכת הקדמות מודל נטייה בשפה טבעית
From Weights to Words: Expressing and Editing Preference Model Inferences in Natural Language
חוקרים פיתחו שיטה חדשה לביטוי הקדמות מודלים בשפה טבעית. השיטה, הנקראת 'משקלות למילים', מאפשרת לגלות ממדים של נטיות בתוך נתונים רב-ממדיים ולתרגמם לשפה טבעית. החוקרים הדגימו את השיטה במספר תחומים, כולל דילמות מוסריות ובחירת סרטים.
תקציר מקורי באנגליתarXiv:2607.16232v1 Announce Type: new Abstract: The growing use of statistical learning algorithms to infer human preferences from high-dimensional choice data runs up against a fundamental challenge: choice alternatives typically differ in many ways simultaneously, so it is generally unclear which factors actually drove an observed decision and should be credited as preferences. Compounding this problem, the opacity of these methods leaves human operators unable to inspect, contest, or correct models when they err. We introduce \emph{weights to words}, a method that takes a dataset of choice problems as input and automatically discovers a collection of domain-relevant preference dimensions, each described in natural language and paired with a vector in the model's representational space.
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית