יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

SharedSAE: מילון אחד לכל מודלי השפה

SharedSAE: One Feature Dictionary Across Language Models
SharedSAE הוא שיטה חדשה המאפשרת שימוש במילון אחד לכל מודלי השפה. השיטה משלבת מילון משותף עם זוגות מקודד-פוענח ספציפיים לכל מודל. SharedSAE שומרת על 96.6% מהשונות המוסברת של SAE מקורי, ומאפשרת העברה של תיאורים לטנטיים בין מודלים.
תקציר מקורי באנגליתarXiv:2609.04344v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) are widely used to interpret language model activations, but SAE training and latent labelling are typically repeated for every model. Here, we show that a single shared SAE can replace a collection of dedicated per-model SAEs. Our method, SharedSAE, combines a shared dictionary with model-specific encoder-decoder pairs. Unlike the closest prior method, which discards activation magnitudes and requires all models at inference, SharedSAE instead normalizes only selection scores, preserving magnitudes, and uses model dropout for single-model inference. We train SharedSAE on four 1B-scale base language models spanning distinct families and tokenizers. Despite sharing its latents across models, SharedSAE retains 96.6%
קרא במקור המקורי