יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

קוונטיזציה של טרנספורמרים מחוברים

Quantizing Looped Transformers: Feedback Exposure and Calibration Blindness
חוקרים זיהו שני מצבי כישלון של קוונטיזציה סטנדרטית. הם מצאו שקוונטיזציה של מודלים מחוברים כמו Huginn-3.5B יכולה לגרום לשגיאות בשלבים מאוחרים יותר. החוקרים הצליחו לשפר את הדיוק על ידי צבירת הסיאן של GPTQ לאורך צעדי החזרה.
תקציר מקורי באנגליתarXiv:2609.30820v1 Announce Type: cross Abstract: Looped transformers reuse weights across recurrence steps, making low-bit quantization especially attractive. We identify two distinct failure modes of standard post-training quantization. On Huginn-3.5B, per-channel INT4 fails primarily at the non-residual loop-entry adapter, while quantizing the residual core is much less damaging. We call this feedback exposure: a quantized layer perturbs the recurrent state without an identity path, and the resulting error is fed back at later steps. Controlled experiments on linear filters and Mamba state-space models show that feedback exposure also occurs outside transformers. Grouped INT4 reveals a separate failure, calibration blindness: our one-step GPTQ baseline builds its Hessian from step-0 act
קרא במקור המקורי