כתבה
arXiv cs.LG ·
קצת בתים, חוק אחד: כיוון ל-W2A4KV2
Few Bits, One Law: Toward W2A4KV2
מציגים את CanonQ, פלטפורמת הסתברות-מודעת-קונטרול-מודעת לאימון ניפוי-כמות-נמוכה של LLM. הפלטפורמה עובדת על קודבוקים-קפאות של טנסורים-שונים, ומאפשרת אימון-שיתוף-פעולה של הרשת לטובת-קונטרול-מודעת.
תקציר מקורי באנגליתarXiv:2610.09202v1 Announce Type: cross Abstract: Extreme low-bit LLM compression is most challenging when weights, activations, and KV caches are quantized together: their distributions differ, and quantization errors interact throughout the network. We introduce CanonQ, a unified quantization-aware training framework that addresses these challenges by separating source canonicalization from task-aware adaptation. Fixed rotations and energy normalization map heterogeneous tensor sources to canonical coordinates, enabling frozen Gaussian-reference codebooks to be reused across layers and models. Joint training then adapts the network to the coupled errors of weight, activation, and cache quantization within a common scalar/vector interface. We bound frozen-codebook transfer error and local
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית