יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

HHR: חיפוש באמצעות קוד חאשי רפרנציאלי לפענוח LLM

HHR: Hierarchical Hash Retrieval for Efficient LLM Generation
HHR: פענוח פענוח LLM באמצעות חיפוש באמצעות קוד חאשי רפרנציאלי, כדי לצמצם את העומס המחשבי. המאמר מציג פרקטיקה חדשה של חיפוש באמצעות קוד חאשי רפרנציאלי, המשלבת תרגום רפרנציאלי ופרויקציה. הפרקטיקה נבחנה במספר תחומים, כולל פענוח LLM, והתוצאות היו טובות. המאמר גם מציג קוד פתוח של הפרקטיקה.
תקציר מקורי באנגליתarXiv:2610.01230v1 Announce Type: new Abstract: Efficient long-context inference is essential for large language models (LLMs), yet it poses a severe computational bottleneck. Hash-based retrieval offers an efficient alternative by encoding queries and keys into binary codes and using Hamming distance for key selection. However, this leads to a critical mismatch between Hamming distance and attention relevance. Query-Key logits depend jointly on directional similarity and feature magnitudes, whereas hash binarization discards magnitude information, causing both false-positive retrieval of low-logit keys and false-negative omission of high-logit keys. To address these failures, we propose Hierarchical Hash Retrieval (HHR), a coarse-to-fine framework that progressively improves retrieval acc
קרא במקור המקורי