יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מניפולציות של קטיעת קומפוננטות ומימדים במנגנוני תיקון תגובה במודלי תרגום

Component and Dimension Sparsity in Transformer Refusal Mechanisms
במאמר זה, חוקרים פורצי דרך חוקרים את מנגנוני תיקון תגובה במודלי תרגום, ומצאו כי הם מורכבים מקומפוננטות ומימדים ספורדיים. המחברים פיתחו כלי למניפולציות של קטיעת קומפוננטות ומימדים, והציגו תוצאות מעניינות.
תקציר מקורי באנגליתarXiv:2610.06903v1 Announce Type: cross Abstract: Activation steering manipulates large language model behavior by intervening on internal activations, but the mechanistic basis of these interventions remains poorly understood. We decompose refusal steering into component-level interventions across four open-weight models, identifying the sparse subsets of attention and MLP components whose steering suffices to reproduce the full behavioral effect. We find that refusal directions concentrate in sparse component mechanisms comprising 28--48\% of upstream components, retaining 88--101\% of steering effectiveness. Within these mechanisms, effective steering further concentrates in approximately 50\% of residual stream dimensions, retaining 85--98\% of the component-mechanism baseline, consist
קרא במקור המקורי