כתבה
arXiv cs.LG ·
פורמט-מודע חיבור לפריתון מהיר FP4
Format-Aware Fusion for Fast FP4 Pretraining
אנו מציגים פורמט-מודע חיבור, שמעבד את כל הקישוריות והפקדות הקישוריות של FP4 כדי להגביר את הביצועים. המאמר עוסק בפיתוח FP4 ובאפשרויות הפיתוח שלו. המאמר כולל גם תיאור של FP4 ושל הפיתוח שלו.
תקציר מקורי באנגליתarXiv:2610.00053v1 Announce Type: new Abstract: Four-bit floating-point (FP4) Tensor Cores accelerate matrix multiplication, but scale computation, operand packing, layout construction, and saved backward state can erase the gain. We present \emph{format-aware fusion}, which co-designs each quantization producer with its scale domain and consumer layout for native \mxfp{}, global \nvfp{}, and cooperative-thread-array-local \nvfp{}. We evaluate Llama-3-family 8B pretraining through 160 billion tokens using bfloat16 output projections and compiled cross entropy. In matched same-accelerator probes, bfloat16 and Transformer Engine \nvfp{} reach 18.8K and 27.6K tokens/s/GPU, while our fastest custom route reaches 37.9K. \mxfp{} with row-gradient stochastic rounding and fixed-sign 32-value Hadam
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית