כתבה
arXiv cs.LG ·
PRQ-KMeans: קידוד טוקנים סמנטיים
PRQ-KMeans: Projection Residual Quantization for Semantic ID Tokenization
PRQ-KMeans הוא אלגוריתם חדש לקידוד טוקנים סמנטיים. הוא משפר את הביצועים של קידוד טוקנים קיימים על ידי הסרת רכיבים גלובליים ושיפור המרכזים. הניסויים הראו שיפורים של עד 7.4% ב-HitRate ו-11.8% ב-MRR.
תקציר מקורי באנגליתarXiv:2608.24207v2 Announce Type: replace Abstract: Semantic identifiers (SIDs) represent entities as hierarchical token sequences for generative retrieval and recommendation. Residual-quantization tokenizers construct these sequences by selecting a codeword at each level and passing a residual to the next. We view this process as progressive commonality removal: each token captures a component shared within its group, while later tokens should model the remaining differences. This view reveals three limitations: a corpus-wide shared component can consume first-level capacity, hard assignment ignores graded similarities to nearby codewords, and full-codeword subtraction can leave variation along the selected-codeword direction in the next residual. We therefore develop our solution in the
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית