כתבה
arXiv cs.AI ·
מעקב אחר מרחב התכונות לצורך זיהוי תכנות לא-מתואמת בעת תקיפה סופרוויזד
Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning
במאמר זה, נציגים שיטה למעקב אחר שינויים במרחב התכונות של רשתות חישוביות בעת תקיפה סופרוויזד. השיטה נועדה לזהות תכנות לא-מתואמת שעלולה להתפתח בעת תקיפה. המחברים הציגו תוצאות מבחן של שיטתם, והראו שהיא יעילה בזיהוי תכנות לא-מתואמת.
תקציר מקורי באנגליתarXiv:2606.07631v2 Announce Type: replace-cross Abstract: Emergent misalignment (EM) occurs when narrow finetuning induces dangerous behavior outside the finetuning task. Detecting this shift through repeated behavioral evaluation is costly, motivating our checkpoint-level monitoring from internal representations. We define a fixed coordinate system from seven alignment-relevant activation directions and use it to track representational drift during LoRA finetuning of four open-source 7-9B language models. Finetuning drift in this space exhibits a dominant axis that explains 78.6% of variance and remains stable across datasets, extraction choices, and parameter-update capacities. Across 468 checkpoints from three EM-relevant held-out datasets, the resulting monitors attain 1.8% FNR, 2.0% F
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית