יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

דיוק לא דיוק: שיטה-בדלנית לבדיקת אבדן נאמנות ב-LLMs מצומצמים

Accuracy is Not Enough: A Divergence-Based Approach to Evaluate Fidelity Loss in Quantized LLMs
במאמר זה, המחברים מציגים שיטה-בדלנית לבדיקת אבדן נאמנות ב-LLMs מצומצמים. הם משתמשים במדדי דחיסה סטטיסטית, כגון דחיסה של ג'נסן-שאנון ודחיסה של טוטל וריאציה, כדי לבדוק את השינויים בתפוצה הפרדקטיבית של ה-LLMs. התוצאות מציגות כי המדדים האלה עולים עם צמצום חזק יותר של ה-LLMs.
תקציר מקורי באנגליתarXiv:2609.07664v2 Announce Type: new Abstract: Deployment of Large Language Models (LLMs) on memory-constrained edge devices relies heavily on aggressive post-training quantization. However, evaluating these models is largely based on zero-shot task accuracy, which depends solely on argmax predictions and is insensitive to changes in the underlying predictive distribution. Consequently, accuracy can exhibit unstable, non-monotonic behavior under progressive quantization, masking substantial fidelity loss relative to the BFloat16 (BF16) uncompressed base model and providing misleading deployment signals. We introduce a distribution-sensitive evaluation framework quantifying information loss in quantized LLMs as the divergence between full-vocabulary predictive distributions at the token de
קרא במקור המקורי