כתבה
arXiv cs.AI ·
The Unequal Influence of Bad Advice: Using Training Data Attribution to Modulate Emergent Misalignment
תקציר מקורי באנגליתarXiv:2609.37914v1 Announce Type: cross Abstract: Fine-tuning large language models on narrow, misaligned tasks can undo their post-training alignment and induce novel misaligned behaviors -- a phenomenon known as \emph{emergent misalignment} (EM). EM has been linked to persona-like representations, where fine-tuning might reduce loss by amplifying a harmful or 'evil' persona. It remains unclear which properties of the training data drive this effect: whether all harmful examples contribute approximately equally to misalignment and whether different models are equally affected by the same fine-tuning examples. In this work, we use training data attribution to quantitatively estimate how much each harmful example contributes to EM. We benchmark the quality of the attribution via retraining
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית