כתבה
arXiv cs.LG ·
LocUS: בקרת הפעלה מכוונת
LocUS: Head Selection and Subspace Projection for Targeted Activation Steering
LocUS היא שיטה לבקרת הפעלה מכוונת במודלים גדולים. היא מאפשרת התערבות בפעילות המודל מבלי לפגוע ביכולות אחרות. LocUS רואה אור בarXiv.
תקציר מקורי באנגליתarXiv:2609.31122v1 Announce Type: cross Abstract: Activation steering is a powerful training-free paradigm for controlling large language models at inference time. However, standard approaches estimate a per-layer steering direction from contrastive data and apply it on the layer's entire representation space, which may couple the intervention to off-target properties present in the contrastive data and degrade unrelated capabilities. To mitigate this issue, we introduce LocUS (Localized Unembedding Steering), a method which grounds activation steering to the model's own output vocabulary subspace. By identifying a property-specific linear subspace within the unembedding matrix, LocUS enforces a geometric constraint that restricts the steering transformation to a specific subspace and at t
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית