כתבה
arXiv cs.CL ·
המרה מובנית לקוונטיזציה בעלות נמוכה של מודלי שפה
Structured Transforms for Low-Overhead Quantization of Language Models
חוקרים הציגו שיטה חדשה לקוונטיזציה של מודלי שפה, המשתמשת בהמרה אורתוגונלית מבוססת DCT. השיטה מאפשרת קוונטיזציה בעלות נמוכה ושומרת על יציבות נומרית. היא נבדקה על מודלים כמו Llama-2 ו-Pythia.
תקציר מקורי באנגליתarXiv:2609.11687v1 Announce Type: new Abstract: We revisit Kashin-decomposition-based weight quantization for large language models and propose an improved algorithm with stronger convergence properties and structured, efficient orthogonal transforms. The method retains the core factorization of each weight into two components -- one with bounded infinity norm and the other with bounded infinity norm after an orthogonal transformation -- but replaces the dense random orthogonal matrix with a sign-randomized Discrete Cosine Transform (DCT), reducing the per-iteration cost from $\mathcal{O}(N^2)$ to $\mathcal{O}(N \log N)$. The proposed greedy algorithm with alternating updates guarantees the four-peak distribution required for stable 2-bit clustering of each factor and admits closed-form in
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית