כתבה
arXiv cs.AI ·
Matryoshka Hash Representations for Model-Aware Compact Semantic Retrieval
תקציר מקורי באנגליתarXiv:2609.07276v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) depends on dense retrieval: each document is stored as a learned vector, and a query is answered by finding its nearest neighbors in that vector space. Keeping one full-precision vector per document is the dominant index cost at corpus scale, so retrieval systems replace each vector with a short code of a few bytes---a step called quantization. Standard quantizers such as product quantization (PQ) pick the code that reconstructs the original vector most closely. A single code is even more useful if it serves several byte budgets at once: when its short prefixes are each directly searchable, a deployment can set its efficiency--quality operating point without re-encoding the corpus. But training all prefi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית