כתבה
arXiv cs.AI ·
ShatterQuant: פריצה לדקדוק זהירות בתקן תצוגה על פלטפורמת מעבד טרנספורמר
ShatterQuant: Breaking Uniform Precision with Block-Wise Mixed-Precision on a Systolic Transformer Hardware Accelerator
ShatterQuant מציגה פלטפורמת מעבד טרנספורמר שמאפשרת דקדוק זהירות בתקן תצוגה. הפלטפורמה משתמשת במעבד טרנספורמר סיסטולי שמאפשר תצוגה מעורבת של 1/2/4/8-bit. ShatterQuant מציגה גם תכונות חדשות של תצוגה מעורבת, כולל תצוגה של 1/2/4/8-bit ופענוח של 1/2/4/8-bit.
תקציר מקורי באנגליתarXiv:2610.00207v1 Announce Type: cross Abstract: Due to limited support for intra-tensor heterogeneous precision in conventional accelerators, neural network quantization remains largely restricted to per-tensor precision assignment. We present ShatterQuant, a hardware-software co-designed framework enabling mixed-precision quantization within each tensor by assigning independent bit-widths to blocks of a weight projection. ShatterQuant couples precision granularity with PE configuration, such that each precision determines an effective block height. We introduce (1) a hardware-aware post-training method that assigns intra-tensor precision based on block-level standard deviation and weight sensitivity; (2) the ShatterQuant Transformer Accelerator supporting 1/2/4/8-bit weight precision, p
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית