כתבה
arXiv cs.CL ·
AlignQuant: קוונטיזציה משולבת ליעילות LLM
AlignQuant: Tile-Aligned Mixed-Precision Quantization for Efficient LLM Generation
AlignQuant היא שיטה לקוונטיזציה משולבת לדגמי LLM. היא משתמשת ביחידות משקל דו-ממדיות על מנת לשפר ביצועים. ניסויים הראו עלייה של עד 2.50 פעמים במהירות יצירה לעומת BF16, תוך שמירה על איכות המודל.
תקציר מקורי באנגליתarXiv:2610.07457v1 Announce Type: cross Abstract: Fine-grained mixed-precision quantization promises efficient large language model inference, but local precision choices can conflict with regular GPU storage and computation units. This precision-boundary mismatch limits the translation of compression into practical acceleration. We introduce AlignQuant, a post-training quantization method that uses GPU-compatible two-dimensional weight tiles as the common unit of precision allocation, compact storage, and execution. This shared partition lets precision follow sensitivity within output channels. Joint prefill/decode calibration scores precision reductions using projection-output perturbations weighted by language-model loss gradients under quantized activations. Phase-normalized scores pri
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית