יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

הנעת רשתות עבור מאפיינים של מקודד אוטומטי דליל

Steering grids for sparse-autoencoder features: when a top-context label names an activation regime rather than a causal axis
חוקרים בדקו את השיטה הסטנדרטית לפרשנות מאפיינים של מקודד אוטומטי דליל. הם הראו שהשיטה הנוכחית מפספסת מידע חשוב. הניסויים בוצעו על מודלים Qwen3-1.7B-Instruct, Gemma-2-2B-it ו-Llama-3.1-8B-Instruct.
תקציר מקורי באנגליתarXiv:2605.03160v2 Announce Type: replace Abstract: The standard protocol for interpreting sparse-autoencoder (SAE) features labels each feature from its top-activating contexts and validates the label by steering that single feature at a typical magnitude. We argue that this inspects one cell of a larger steering grid, steering condition (single feature, joint feature set, matched random direction) crossed with steering coefficient, and show that other cells carry information that changes the label. On Qwen3-1.7B-Instruct and Gemma-2-2B-it, with the matched-geometry control extended to Llama-3.1-8B-Instruct: (1) features labelled AI self-disclaimer from their top contexts switch to a second surface form under steering, a contemplative voice on Qwen, a collective we-voice on Gemma, so the
קרא במקור המקורי