יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

TR-PTQ: קוונטיזציה של טרנספורמרים

TR-PTQ: High-Accuracy Integer-Only Transformer Post Training Quantization via Taylor Region Reformulation
TR-PTQ הוא שיטה חדשה לקוונטיזציה של טרנספורמרים, המאפשרת יישום יעיל עם דיוק גבוה. השיטה מתמודדת עם בעיות הדיוק הנובעות משכבות לא ליניאריות.
תקציר מקורי באנגליתarXiv:2610.09969v1 Announce Type: new Abstract: Post-training quantization (PTQ) enables efficient deployment, yet transformer architectures remain challenging to quantize due to nonlinear layers. While existing methods attribute accuracy loss to insufficient numerical precision, often necessitating floating-point fallbacks, we demonstrate that degradation is actually driven by specific structural error sources. We find that learned scale parameters in normalization layers and compounded approximations in GELU are the primary error contributors, whereas SoftMax remains inherently robust to aggressive quantization. To address these bottlenecks, we introduce TR-PTQ, a unified integer-only formulation using shared Taylor Region (TR) exponential and logarithm primitives. This approach allows c
קרא במקור המקורי