כתבה
arXiv cs.LG ·
Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders
תקציר מקורי באנגליתarXiv:2605.07922v3 Announce Type: replace Abstract: Learning hierarchical features in Sparse Autoencoders (SAEs) is essential for capturing the structured nature of real-world data and mitigating issues like feature absorption or splitting. Existing works attempt to identify hierarchical relationships within independent feature sets by relying on activation coverage, the assumption that child feature should only activate when its parent feature activates. However, we demonstrate that this condition alone is insufficient; that is, it often produces false positives where parent and child concepts are semantically unrelated. To address this, we introduce a novel reconstruction condition that enforces a deeper functional link between hierarchical levels. By combining both activation and recons
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית