כתבה
arXiv cs.LG ·
Do Sparse Autoencoders Learn Meaningful Concept Hierarchies?
תקציר מקורי באנגליתarXiv:2606.22994v2 Announce Type: replace Abstract: Sparse autoencoders (SAEs) have become an important tool for unsupervised concept discovery in large models. To make the resulting feature spaces more interpretable and manageable, recent approaches have begun imposing hierarchical structure, either explicitly or as an implicit effect of training constraints, yet rigorous comparison remains difficult. There are no agreed-upon requirements for what a meaningful feature hierarchy should satisfy, and evaluation has largely relied on qualitative illustrations with fragmented quantitative protocols. To address this, we derive a set of key requirements for generalization/specialization hierarchies in unsupervised concept discovery, drawing on semantic net and taxonomy research alongside recent
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית