יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

כל אבלציה היא דוז': נגדי-משקלים והדמיון של תיקון עצמי

Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair
אבלציה במודלי שפה מגלה השפעה של נגדי-משקל. נראה שהמודלים תוקנים את עצמם, אך זה רק תגובה של נגדי-משקל שמבצע את תפקידו הרגיל. המאמר עוסק בזיהוי נגדי-משקלים במודלי GPT-2 ובמודלים אחרים.
תקציר מקורי באנגליתarXiv:2610.02173v1 Announce Type: new Abstract: Ablate a component of a language model, and other components often appear to adjust and compensate. This phenomenon, termed self-repair, has been observed repeatedly, but its mechanism remains unclear. The most systematic study to date concluded that self-repair is noisy and unlikely to have a single explanation. We argue that it has one: a gain already present before any ablation. Any intervention on a causally important component can be viewed as a point on a coordinate axis $\lambda$, the signed strength of a counterfactual contrast. Hence, conventional ablation methods are uncalibrated points on this axis. We show that the causal repair response for a fine-grained unit $r$ is governed by an affine law, $E_r(\lambda)=\mathrm{own}_r+\gamma_
קרא במקור המקורי