כתבה
arXiv cs.LG ·
קוונטיזציה על ידי Drift: קוונטיזציה לא-מאותגגת לאחר האימון ללא תוויות למעבדי טקסט
Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders
קוונטיזציה לא-מאותגגת לאחר האימון ללא תוויות למעבדי טקסט. פיתוח חדשני לקוונטיזציה של מעבדי טקסט.
תקציר מקורי באנגליתarXiv:2610.09227v1 Announce Type: cross Abstract: Mixed-precision post-training quantization needs a per-module sensitivity signal; for a text embedder the obvious one -- the retrieval quality a module costs when quantized -- needs relevance labels that deployments rarely have. We measure a label-free substitute: quantization-induced representation drift, obtained by quantizing one module, re-encoding the corpus, and recording how far the output embeddings moved from their full-precision positions. What is specific is the observable: the deployed output representation a dense retriever ranks with. Across five development embedders, configuration-level drift orders sampled mixed-precision plans against held-out retrieval quality at a macro Spearman of 0.911, the sensitivity transports acros
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית