יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

עברי: עבר את ההתאמה השטוחה: איך שיטות פוסט-אימון מגדירות תפקידי סירוב ואמינות שליטה

Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness
שיטות פוסט-אימון משפיעות על חישוב הסירוב של מודלי שפה. נחקרו שלוש שיטות: טיפול סופרוויזד, טיפול סופרוויזד עם רציונליזציה, ואופטימיזציה של קונפורמציה. נמצא כי שיטת הרציונליזציה יוצרת סירוב חדשני, ושאין שיטה שמגיעה לשלושת התכונות הרצויות: סירוב שאינו קונצנטרט, תועלת בבטיחות שאינה עולה בתועלת כללית, וביצועי בטיחות שניתנים לתיקון דרך תיקונים קטנים.
תקציר מקורי באנגליתarXiv:2609.03887v1 Announce Type: new Abstract: How do the methods used to train language models to refuse harmful requests shape how that refusal actually works inside the model? We compare three post-training methods - supervised fine-tuning, reasoning-augmented fine-tuning (training on reasoning chains that justify a safety decision), and preference optimization (ORPO) - across three architecturally distinct models (Llama-3.1-8B, Gemma-2-9B, Qwen3-8B). We find that training method, not just data, reshapes how refusal is computed internally: reasoning-augmented training consistently produces a distinct kind of refusal computation, visible across all three models, while architecture independently shapes internal structure and how reliably refusal can be steered. Most importantly, no metho
קרא במקור המקורי