יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

השפעה לא שווה של עצות גרועות: שימוש באטריביוציה של נתוני האימון למודולציה של התאמה נבחרת

The Unequal Influence of Bad Advice: Using Training Data Attribution to Modulate Emergent Misalignment
אימון מודלי שפה גדולים על תפקודים צרים ולא מתואמים יכול לגרום להתנהגויות לא מתואמות חדשות. המחקר חקר את השפעת נתוני האימון על תוצאות האימון. התוצאות הראו שכל המודלים המבודדים הפכו ללא מתואמים כאשר נאמן על אותו הסט הנתונים, והציונים הטובים ביותר של האטריביוציה היו כאשר נאמן על נתוני האימון של המודל שממנו נקבעו.
תקציר מקורי באנגליתarXiv:2609.37914v1 Announce Type: new Abstract: Fine-tuning large language models on narrow, misaligned tasks can undo their post-training alignment and induce novel misaligned behaviors -- a phenomenon known as \emph{emergent misalignment} (EM). EM has been linked to persona-like representations, where fine-tuning might reduce loss by amplifying a harmful or 'evil' persona. It remains unclear which properties of the training data drive this effect: whether all harmful examples contribute approximately equally to misalignment and whether different models are equally affected by the same fine-tuning examples. In this work, we use training data attribution to quantitatively estimate how much each harmful example contributes to EM. We benchmark the quality of the attribution via retraining --
קרא במקור המקורי