כתבה
arXiv cs.AI ·
OSFP4: אופטימיזציה משולבת לקידום NVFP4
OSFP4: Joint Optimization of Diagonal Smoothing and Block Scales for NVFP4 Quantization
OSFP4 הוא שיטה חדשה לאופטימיזציה של קידום NVFP4, המאפשרת שמירה על דיוק וביצועים גבוהים. השיטה משתמשת במטריצת חלקה אלכסונית ואופטימיזציה משולבת של קנה מידה. ניסויים הראו ש-OSFP4 משיג את הדיוק הממוצע הגבוה ביותר בהשוואה לשיטות אחרות.
תקציר מקורי באנגליתarXiv:2610.08231v1 Announce Type: new Abstract: NVFP4 is an attractive datatype for large language model (LLM) inference, offering compact storage and native tensor-core acceleration. However, preserving accuracy using NVFP4 requires careful quantization. In this work we develop a novel quantization scheme called Optimized Smoothing and Scaling for NVFP4 (OSFP4). For each linear projection it uses a diagonal smoothing matrix whose entries are optimized to minimize the squared matrix-product quantization error under NVFP4, taking into account the rounding procedure that is used (either round-to-nearest, or GPTQ-style successive interference cancellation). This requires performing joint optimization on the smoothing entries as well as the block scales, which is facilitated by analyzing a mul
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית