כתבה
arXiv cs.LG ·
השפעת ההפרעה משקפת את ההתנהגות היסודית של המודל
Steering Interference Reflects the Model's Defaults, Not the Behavior Directions
מחקר חדש מצביע על כך שההשפעה של ההפרעה במודלי שפה משקפת את ההתנהגות היסודית של המודל, ולא את ההתנהגות הרצויה. המחקר בדק 24 התנהגויות שונות ו-10 מודלים, ומצא שההשפעה היא תוצאה של המודל עצמו, ולא של ההתנהגות הרצויה. התוצאות הללו עלולות להשפיע על הדרך בה אנו מתכננים ומאמנים מודלי שפה.
תקציר מקורי באנגליתarXiv:2609.06951v1 Announce Type: new Abstract: Activation steering promises modular control of language model behavior: a behavior such as politeness corresponds to a direction in a model's activations, and adding that direction while it generates should switch the behavior on and leave everything else alone. It does not. We ask what decides which other behaviors move, and by how much, and find that it is the model rather than the behavior being steered. A steer relaxes the model toward a small set of behaviors it already favors, chiefly refusal, sycophancy, and poeticism, and that set is much the same whatever is steered. Three results across 24 behaviors and ten instruction-tuned models support this, every effect read off the generated text by a language-model judge rather than off a pr
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית