יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

אינדוקציה אינה מספיקה: אימות סיבתי של תכונות SAE רב-לשוניות

Recurrence Is Not Enough: Causally Validating Multilingual SAE Translation Features in Gemma 2 and 3
חוקרים בדקו את תפקידן של תכונות SAE רב-לשוניות במודלים Gemma 2 ו-3. הם מצאו כי רוב התכונות שנמצאו במספר שפות לא השפיעו על התנהגות התרגום. רק תכונה אחת השפיעה באופן עקבי על תוצאות התרגום.
תקציר מקורי באנגליתarXiv:2609.04808v1 Announce Type: new Abstract: Sparse autoencoder (SAE) features are increasingly used to explain and steer language-model behavior, but it remains unclear whether a feature found in one language context plays the same causal role when processing prompts in another language. We study this question using translation-initiation features (Wu et al., 2026). We reproduce the SAE feature discovery method from Wu et al. in Gemma 2 and extend it to multilingual settings that vary prompt language, source language, and target language. We then test whether features that recur across settings affect translation behavior by amplifying or ablating their activations during inference. We also examine whether the method can be applied to Gemma 3. In both models, we observe an identical fi
קרא במקור המקורי