יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

איך לבטוח בכלים לפירוש מודלי שפה

Calibrating Interpretability Instruments Before Trusting Their Verdicts
חוקרים מציגים שיטה לבדיקת כלים לפירוש מודלי שפה. הם מדגימים שישויות פוטנציאליות בכלים אלו, ומציעים פרוטוקול לווידוא אמינותם. המחקר בוחן ארבעה מודלים, כולל LLaMA.
תקציר מקורי באנגליתarXiv:2609.14754v1 Announce Type: cross Abstract: Causal claims about large language model (LLM) internals rest on measurements. Those might include a projection, a cosine, an ablation delta, or an interchange patch among others. These measurements fail in specific, diagnosable ways that return a plausible number instead of an error, so a broken instrument can easily read as a finding. A covariance-matched null can saturate until every direction looks typical, a per-head attribution can overshoot the true residual write threefold on reordered-normalization architectures, an interchange patch can go sign-chaotic because its outcome is pinned at a ceiling, or a read-from verdict can be an artifact of measuring past the layer where the model already decided. This note documents six such failu
קרא במקור המקורי