כתבה
arXiv cs.AI ·
המפענח: חיזוק הביטוי של הוראות בטיחות באמצעות אופטימיזציה ספקטרלית
The Safety Operator: Modulating the Expression of Safety Instructions via Spectral Optimization
במאמר זה, חוקרים חוקרים את האפשרות לחיזוק הביטוי של הוראות בטיחות באמצעות אופטימיזציה ספקטרלית. הם מציעים פונקציית חיסול של בטיחות שמאפשרת לשלוט במידת ההשפעה של הוראות הבטיחות על הפעילות של המודל.
תקציר מקורי באנגליתarXiv:2609.36434v1 Announce Type: new Abstract: Context tokens in a transformer-based language model can be absorbed into the model's weights as a multiplicative operator. We study this operator in the setting of safety instructions and show that influencing its dominant eigenvalue modulates how strongly the instruction shapes generation. We derive a Contrastive Safety Loss with a suppression weight that controls the tradeoff between emphasizing the safety instruction on harmful queries while suppressing it on harmless queries. Varying the suppression weight maps a relationship between the attack success and the over-refusal rates, supporting the hypothesis that the operator's eigenvalue acts as a continuous dial for the instruction's influence. Moreover, this relationship holds relatively
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית