כתבה
arXiv cs.CL ·
אי-התאמה עם אישיות: חשבון Big Five של אי-התאמה מתעוררת
Misalignment Has a Personality: A Big Five Account of Emergent Misalignment
חוקרים גילו כי אי-התאמה במודלי שפה קשורה למאפייני אישיות. המחקר מראה כי מודלים שעברו עידון עדין על נתונים בעייתיים מפגינים אישיות שונה, עם ירידה בהסכמה ומודעות עצמית ועליה בחריצות ונוירוטית.
תקציר מקורי באנגליתarXiv:2607.26389v1 Announce Type: new Abstract: Fine-tuning a language model on data containing a narrow flaw, such as insecure code or incorrect mathematical answers, can cause broad misalignment through a mechanism that remains debated. We provide an interpretable account: in the models and corpora we study, misalignment behaves like a shift in personality. Prior work extracts activation directions for character traits from a single binary contrast, which can separate or steer behavior without establishing a calibrated scale. We instead extract personality vectors for the Big Five using a graded, three-level intervention and validate them on two open-weight models. The three levels are linearly ordered, with Cohen's d values of up to 6.2; the vectors transfer zero-shot and trait-specific
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית