כתבה
arXiv cs.CL ·
SeDeM: דחיסה סלקטיבית לשאילתות ארוכות
SeDeM: Selective Decompression of Hidden-State Memories for Long-Context Question Answering
SeDeM הוא כלי חדש לדחיסה סלקטיבית של זיכרון המאפשר עיבוד יעיל יותר של שאילתות ארוכות. הוא משמש לדחיסת זיכרון המאפשרת למודלים לעבד שאילתות ארוכות יותר. SeDeM משיג תוצאות טובות יותר משיטות דחיסה אחרות במספר מבחנים.
תקציר מקורי באנגליתarXiv:2608.00311v2 Announce Type: replace Abstract: Long-context inference with large language models (LLMs) is costly: self-attention during prefill scales quadratically with sequence length, and the key-value (KV) cache grows with the number of processed tokens. Larger context windows also do not ensure reliable evidence use. Context compression reduces this cost, but many soft-compression methods use LLMs as compressors and rely on compact memory tokens both to preserve information and to condition the decoder. We propose SeDeM, a selective decompression framework that decouples compact memory storage from decoder conditioning. SeDeM stores context as compact hidden-state memory blocks, selects query-relevant blocks, and decompresses only the selected blocks for decoder conditioning. Th
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית