כתבה
arXiv cs.CL ·
SAEExplainer: פירוש תכונות SAE עם אופטימיזציה מונחית הפעלה
SAEExplainer: Interpreting SAE Features with Activation-Guided Preference Optimization
SAEExplainer הוא כלי לפירוש תכונות SAE. הוא משתמש באופטימיזציה מונחית הפעלה כדי לשפר את ההסברים. הכלי מוכיח עליונות על בסיסי השוואה קיימים.
תקציר מקורי באנגליתarXiv:2606.08496v2 Announce Type: replace Abstract: Although Sparse Autoencoders (SAEs) have mitigated the opacity of large language models (LLMs) by decomposing dense representations into sparse features, explaining these features still remains a central challenge. Current explanation methods, however, typically operate within an open-loop paradigm, failing to leverage mechanistic feedback for further refinement. In this paper, we propose SAEExplainer, a training framework that utilizes activation scores as an objective reward signal to train the model for self-correction and iterative bootstrapping. By iteratively verifying and correcting foundational explanations through a two-round optimization process, SAEExplainer achieves continuous improvement in its explanatory capabilities. This
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית