יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

אחסון המקורב: זיכרון קצר וניתן לשימוש מחדש ל-LLM

Cache the Encoder Within:Compact, Reusable Memory across LLM Queries
מערכת אחסון קצרה וניתנת לשימוש מחדש לביצועי LLM, כולל שמות מודלים/כלים/חברות. המאמר עוסק באחסון מקורב של מודל LLM, כולל גמיני ו-GPT-5.
תקציר מקורי באנגליתarXiv:2610.10058v1 Announce Type: new Abstract: Repeated queries over shared documents incur redundant encoding, while caching model states introduces persistent storage costs. Building on CoMem's intermediate-state interface, EncBank treats a pretrained LLM's lower layers as a reusable document encoder and compactly stores their outputs for an adapted upper-layer reader. A self-distilled suffix adapter is shared across storage precisions within each backbone, without quantization-specific retraining. Across five benchmark suites on three Qwen backbones spanning different sizes and full-attention and hybrid architectures, 4-bit storage keeps each reported benchmark aggregate within one score point of native-precision EncBank. In a fixed Qwen3-8B workload, it retains 28.1% of the native-pre
קרא במקור המקורי