יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

SharedSAE: ספריית תצורה אחת למודלי שפה

SharedSAE: One Feature Dictionary Across Language Models
SharedSAE מציעה ספריית תצורה אחת שמספקת הסברה למודלי שפה שונים. המתודה, שפועלת על ידי ספריית תצורה משותפת וזוגות קודר-מקדם מודל-נפרדים, יוצרת הסברה טובה יותר ומשפרת את הסברה של מודלי שפה שונים.
תקציר מקורי באנגליתarXiv:2609.04344v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) are widely used to interpret language model activations, but SAE training and latent labelling are typically repeated for every model. Here, we show that a single shared SAE can replace a collection of dedicated per-model SAEs. Our method, SharedSAE, combines a shared dictionary with model-specific encoder-decoder pairs. Unlike the closest prior method, which discards activation magnitudes and requires all models at inference, SharedSAE instead normalizes only selection scores, preserving magnitudes, and uses model dropout for single-model inference. We train SharedSAE on four 1B-scale base language models spanning distinct families and tokenizers. Despite sharing its latents across models, SharedSAE retains 96.6% o
קרא במקור המקורי