כתבה
arXiv cs.AI ·
רפרמטריזציה של softmax להקטנת קומפרציה של ראשי-מילה
Softmax Reparameterization for Output-Head Quantization
אורחות-פעולה חדשות של softmax להקטנת קומפרציה של ראשי-מילה. ניתן להקטין את ההשפעה של קומפרציה על תוצאות המודל. השיטה נבחנה על ידי ניסויים על מודלי XGLM, Phi, BLOOM ו-BLOOMZ.
תקציר מקורי באנגליתarXiv:2609.31291v2 Announce Type: cross Abstract: Large vocabularies make output heads a substantial inference cost in small language models. We introduce softmax reparameterization, a post-training method that searches over functionally equivalent output heads before quantization. The method subtracts a scalar multiple of the vocabulary-row mean from every output row and selects the coefficient by validation KL. For linear-softmax heads, these shifts preserve full-precision predictions exactly and require no decoder retraining; a rank-one correction extends the construction to nonlinear logit paths. Across seven output heads and three quantizers, W4 gains are largest where baseline quantization substantially distorts predictions: test KL falls by 93% on XGLM under RTN and by 73--77% on Ph
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית