כתבה
arXiv cs.CL ·
קוונטיזציה ללא תוויות
Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders
חוקרים פיתחו שיטה לקוונטיזציה ללא תוויות למשובצים טקסט. השיטה מודדת את השינוי בייצוג הטקסט לאחר קוונטיזציה, ומשתמשת בו כאות לרגישות. השיטה הוכחה כיעילה במספר מודלים.
תקציר מקורי באנגליתarXiv:2610.09227v1 Announce Type: cross Abstract: Mixed-precision post-training quantization needs a per-module sensitivity signal; for a text embedder the obvious one -- the retrieval quality a module costs when quantized -- needs relevance labels that deployments rarely have. We measure a label-free substitute: quantization-induced representation drift, obtained by quantizing one module, re-encoding the corpus, and recording how far the output embeddings moved from their full-precision positions. What is specific is the observable: the deployed output representation a dense retriever ranks with. Across five development embedders, configuration-level drift orders sampled mixed-precision plans against held-out retrieval quality at a macro Spearman of 0.911, the sensitivity transports acros
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית