כתבה
arXiv cs.AI ·
מניעת תסביך נביל: תיקון סבילות כפייתי
Correct, Don't Delete: Mitigating Emergent Misalignment with Corrective Supervision
במאמר זה, המחברים חוקרים את השפעת תיקון סבילות כפייתי על תסביך נביל במודלי שפה. הם מציעים שיטה לתיקון סבילות כפייתי, שמטרתה למנוע תסביך נביל במודלי שפה. השיטה נבדקה במודל Qwen2.5-14B-Instruct והתוצאות היו מוצלחות.
תקציר מקורי באנגליתarXiv:2609.37624v1 Announce Type: cross Abstract: Fine-tuning a language model on a narrow set of harmful demonstrations, such as bad medical advice, can make it broadly misaligned on unrelated questions, a phenomenon known as emergent misalignment (EM). The usual defense is to find the offending rows and delete them, but a row locator failed our held-out test and deleting rows helps less than expected. We ask a different question: given a fixed set of poisoned rows, is it better to correct them than to remove them? We fine-tune Qwen2.5-14B-Instruct on a mixture of bad medical advice and benign chat data, select a quarter of the poison rows in advance, and either delete them or replace each with a corrected answer to the same prompt, keeping everything else the same. Replacing the rows cut
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית