יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

דומה לקוסינוס אינו ראיה

Cosine Similarity Is Not Evidence: Measuring the Noise Floor of Interpretability Transfer Under Quantization
חוקרים בדקו את השפעת הקוונטיזציה על הבנת מודלים. הם מצאו שדומה לקוסינוס אינו מספיק כדי להוכיח שמודלים שורדים שינויים. הם השתמשו במודל Qwen2.5-1.5B-Instruct כדי לבדוק את התוצאות.
תקציר מקורי באנגליתarXiv:2609.30275v1 Announce Type: new Abstract: A statistic reported without the quantity needed to interpret it is not evidence. We develop that thesis for a concrete practice in AI safety. Interpretability artifacts are calibrated on full-precision weights, deployed on quantized ones, and certified as surviving the change by scale-invariant statistics (cosine similarity, correlation, AUROC) that are reported without their noise floor. For the difference-in-means direction estimator, the split-half floor is governed by one dimensionless number, $\kappa = n\rho^2/d$. The closed form $\mathbb{E}[\cos] \approx (1+4/\kappa)^{-1}$ is classical; the missing input is the class separation $\rho$, which we measure on real activations; no compression-transfer study we know of reports it. On Qwen2.5
קרא במקור המקורי