כתבה
arXiv cs.CL ·
עברי: עבר את ההתאמה השטוחה: איך שיטות פוסט-אימון מגדירות תפקידי סירוב ואמינות שליטה
Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness
שיטות פוסט-אימון משפיעות על חישוב הסירוב של מודלי שפה. נחקרו שלוש שיטות: טיפול סופרוויזד, טיפול סופרוויזד עם רציונליזציה, ואופטימיזציה של קונפורמציה. נמצא כי שיטת הרציונליזציה יוצרת סירוב חדשני, ושאין שיטה שמגיעה לשלושת התכונות הרצויות: סירוב שאינו קונצנטרט, תועלת בבטיחות שאינה עולה בתועלת כללית, וביצועי בטיחות שניתנים לתיקון דרך תיקונים קטנים.
תקציר מקורי באנגליתarXiv:2609.03887v1 Announce Type: new Abstract: How do the methods used to train language models to refuse harmful requests shape how that refusal actually works inside the model? We compare three post-training methods - supervised fine-tuning, reasoning-augmented fine-tuning (training on reasoning chains that justify a safety decision), and preference optimization (ORPO) - across three architecturally distinct models (Llama-3.1-8B, Gemma-2-9B, Qwen3-8B). We find that training method, not just data, reshapes how refusal is computed internally: reasoning-augmented training consistently produces a distinct kind of refusal computation, visible across all three models, while architecture independently shapes internal structure and how reliably refusal can be steered. Most importantly, no metho
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית