יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

מה גובה 0.6? קרשים, תקרות וראשי חדר באינטרפרטביליות

How High Is 0.6? Floors, Ceilings, and Headroom in Interpretability Probing
במאמר זה, המחברים חוקרים את המשמעות של ציון R^2 של 0.6 באינטרפרטביליות. הם מציעים לקרוא את הציון נגד שני נקודות תצפית: קרש ותקרה. המחברים מבודקים את הצעתם על דגמי תרגומי מטען, שמחשבים את ההטרוגניות הביולוגית. הם גם מחקרים את המודלים האמיתיים.
תקציר מקורי באנגליתarXiv:2610.08544v1 Announce Type: new Abstract: Probes are the workhorse of interpretability. If a model's hidden states predict a variable, the model is said to represent it. But a probe score has no fixed meaning. An $R^2$ of 0.6 may only reflect what the input already gives away, and the same score can mean different things on different data. We propose reading every probe score against two reference points: a floor, what a declared set of simple inputs already predicts, and a ceiling, what the full input can predict. The gap between them, the headroom, is the range in which a probe can show that a model computes something beyond the simple inputs. We prove that headroom vanishes in two ways: the target stops depending on a hidden variable the model must infer, or the input stops reveal
קרא במקור המקורי