יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

RDQ: קיצור תפוצה של תוצאות נוזליות למודלי שפה גדולים

RDQ: Residual Distribution Quantization for Large Language Models
RDQ היא תפיסה חדשה של קיצור תפוצה של תוצאות נוזליות למודלי שפה גדולים. היא מציעה פתרון לבעיית הדחיסה של מודלי שפה גדולים, ומציעה תוצאות טובות יותר בהשוואה למודלים אחרים.
תקציר מקורי באנגליתarXiv:2607.10137v3 Announce Type: replace Abstract: Post-training quantization (PTQ) of large language models degrades sharply below 4-bit precision. We identify the root cause as residual stream distributional drift: quantization noise injected at each transformer layer accumulates in the shared residual representation, causing KL divergence from the FP16 baseline to grow super-linearly with depth (Pearson r=0.999 with log-perplexity, p<0.001, confirmed across all tested methods and bit-widths). We discover that 84% of LLaMA-3-8B layers exhibit non-Gaussian residual distributions (KS test, p<=0.05), and that per-layer residual stream variance grows 6,548x across depth. We propose RDQ (Residual Distribution Quantization), a PTQ framework whose central contribution is Cascaded Error Compens
קרא במקור המקורי