יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

SAEExplainer: הסבר תכונות SAE

SAEExplainer: Interpreting SAE Features with Activation-Guided Preference Optimization
SAEExplainer הוא כלי להסבר תכונות Sparse Autoencoders (SAEs). הוא משתמש באותות הפעלה כאותות רווח כדי לאמ� את המודל לתיקון עצמי ואימון עצמי. SAEExplainer משפר את היכולות ההסבריות שלו באמצעות תהליך אופטימיזציה כפול.
תקציר מקורי באנגליתarXiv:2606.08496v2 Announce Type: replace-cross Abstract: Although Sparse Autoencoders (SAEs) have mitigated the opacity of large language models (LLMs) by decomposing dense representations into sparse features, explaining these features still remains a central challenge. Current explanation methods, however, typically operate within an open-loop paradigm, failing to leverage mechanistic feedback for further refinement. In this paper, we propose SAEExplainer, a training framework that utilizes activation scores as an objective reward signal to train the model for self-correction and iterative bootstrapping. By iteratively verifying and correcting foundational explanations through a two-round optimization process, SAEExplainer achieves continuous improvement in its explanatory capabilities.
קרא במקור המקורי