כתבה
arXiv cs.LG ·
Capsule Lens: מיקום ועקיבה של גאומטריה של רעיונות בייצוגי מודל
Capsule Lens: Locating and Tracking Concept Geometry in Model Representations
מערכת למיקום ועקיבה של גאומטריה של רעיונות בייצוגי מודל. המערכת נועדה לסייע בהבנת כיצד רעיונות נקודסים בייצוגי מודל.
תקציר מקורי באנגליתarXiv:2609.05575v1 Announce Type: new Abstract: Understanding how concepts are encoded in the internal representations of machine learning models is a central problem in mechanistic interpretability, essential both for the science of deep learning and for the trustworthy deployment of increasingly capable models. Existing approaches to interpret model representations mainly map representations onto more interpretable spaces and do not directly characterize how concepts occupy representation space; various hypotheses have been proposed, but often lack of rigorous validation and largely focus on static representations. In this work, we introduce Capsule Lens, a framework that matches the region a concept occupies with a simple, trackable geometric form, a capsule, defined by several interpre
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית