כתבה
arXiv cs.CL ·
LocUS: ניהול פעילות והצגת תת-מרחב להפעלה מופנת
LocUS: Head Selection and Subspace Projection for Targeted Activation Steering
LocUS היא שיטה שמטרתה לשלוט במודלי שפה גדולים בזמן הריצה, ולמנוע התערבות בתכונות שאינן רלוונטיות.
תקציר מקורי באנגליתarXiv:2609.31122v1 Announce Type: new Abstract: Activation steering is a powerful training-free paradigm for controlling large language models at inference time. However, standard approaches estimate a per-layer steering direction from contrastive data and apply it on the layer's entire representation space, which may couple the intervention to off-target properties present in the contrastive data and degrade unrelated capabilities. To mitigate this issue, we introduce LocUS (Localized Unembedding Steering), a method which grounds activation steering to the model's own output vocabulary subspace. By identifying a property-specific linear subspace within the unembedding matrix, LocUS enforces a geometric constraint that restricts the steering transformation to a specific subspace and at the
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית