יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

חתימות של שליטות במרחב הפעילות של דגמי שפה

Signatures of Steerability in Activation Space of Language Models
במאמר זה, חוקרים חוקרים את יעילות שליטה בדגמי שפה על ידי נציגי תחושה נגדית. הם טוענים שמטרייה-התלויה בדטה-סט זו, ומציעים סטטיסטיקות פשוטות כדי לזהות כאשר שליטה עשויה להיות יעילה.
תקציר מקורי באנגליתarXiv:2609.14151v1 Announce Type: cross Abstract: Steering language models using a set of contrastive representations has been a canonical and computationally efficient method for controlling model behavior. Despite this success in controlling certain model behaviors, the effectiveness of activation steering varies markedly across concepts; the generalization properties of steering vectors are often considered a function of the dataset used to construct them. We make this dataset-dependence claim more rigorous and show that simple separation metrics strongly correlate with the downstream steerability of language models across diverse settings, even after controlling for layers and dataset effects. Beyond prediction, we provide evidence from a synthetic superposition experiment that separat
קרא במקור המקורי