כתבה
arXiv cs.LG ·
תיאור טכני: חידוש באלוקציה של זיכרון לצורך קיצור רשת
Layerwise Error Attribution for Fast and Robust Mixed-Precision Post-Training Quantization
המאמר עוסק בפיתוח שיטה חדשה לאלוקציה של זיכרון ברשתות עבור קיצור רשת. השיטה נותנת תיאור טכני פרובביסטי של טעות הקיצור, ומאפשרת אלוקציה יעילה ורבועה. השיטה נבדקה במספר תחומים, כולל קיצור רשת וקיצור רשת עבור דיפוזיה.
תקציר מקורי באנגליתarXiv:2610.09877v1 Announce Type: new Abstract: Mixed-precision post-training quantization is a network compression method that assigns bits layer by layer, under a global memory budget using a small calibration set. The main difficulties are to overcome the combinatorial nature of the allocation problem and to manage the sensitivity to small, potentially corrupted databases. Hence, an efficient allocation method should be fast to compute and preserve model quality when calibration data are corrupted. To design such a method, we derive a layerwise probabilistic analysis of the quantization error that separates propagated error from the local perturbation introduced at a given layer. We use this local term to build a separable score for a simple allocation algorithm, that requires no extern
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית