כתבה
arXiv cs.AI ·
מה גובה 0.6? קרשים, תקרות וראשי חדר בבדיקת פרובינג
How High Is 0.6? Floors, Ceilings, and Headroom in Interpretability Probing
במאמר זה, המחברים מציגים חידוש בבדיקת פרובינג, שהיא כלי חשוב בבדיקת פרובינג. הם מציעים לקרוא את התוצאות של הבדיקה בהשוואה לשני נקודות תצפית: קרש ותקרה. הקרש הוא התוצאה שהמודל יכול להשיג עם נתונים פשוטים, והתקרה היא התוצאה שהמודל יכול להשיג עם נתונים מלאים. הפער בין השניים, הראשי חדר, הוא הטווח בו הבדיקה יכולה להראות שהמודל מחשב דברים שלא ניתנים לחישוב.
תקציר מקורי באנגליתarXiv:2610.08544v1 Announce Type: cross Abstract: Probes are the workhorse of interpretability. If a model's hidden states predict a variable, the model is said to represent it. But a probe score has no fixed meaning. An $R^2$ of 0.6 may only reflect what the input already gives away, and the same score can mean different things on different data. We propose reading every probe score against two reference points: a floor, what a declared set of simple inputs already predicts, and a ceiling, what the full input can predict. The gap between them, the headroom, is the range in which a probe can show that a model computes something beyond the simple inputs. We prove that headroom vanishes in two ways: the target stops depending on a hidden variable the model must infer, or the input stops reve
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית