יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

השוואה שיטתית של שיטות פרשנות רב-לשוניות

A Systematic Comparison of Multilingual Interpretability Methods Reveals Anisotropy-Driven Failures
חוקרים השווים ארבע שיטות לפרשנות רב-לשונית, ומצאו ש-ILO היא השיטה היעילה ביותר. המחקר בדק 21 מודלים בגדלים שונים, ומצא קורלציה גבוהה בין ILO לביצועים קרוס-לינגואליים.
תקציר מקורי באנגליתarXiv:2609.04819v1 Announce Type: new Abstract: Multilingual language models develop shared cross-lingual representations, and various interpretability methods claim to quantify this sharing. These methods have been developed largely in isolation, and when they disagree, it is unclear whether the disagreement reflects a property of the model or an artifact of the measurement. We compare four sharing metrics (CKA, ANC, GMM dominance per token, and ILO) across 21 base models from five families (125M-14B parameters) and correlate each with cross-lingual transfer on five downstream tasks. We find that the metrics differ in their quantification of cross-lingual sharing in these models and suggest that the disagreement traces to anisotropy, the tendency of representations to cluster in a narrow
קרא במקור המקורי