יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

ATTUNER: שימוש מחדש של קאש KV בלי תקיפה מחדש

ATTUNER: Recomputation-Free KV Cache Reuse via Query-Side Adaptation
ATTUNER מאפשר שימוש מחדש של קאש KV בלי תקיפה מחדש, תוך שימוש בשיטת עיבוד קצרה. השיטה משמשת לשיפור ביצועי מודלי LLM.
תקציר מקורי באנגליתarXiv:2609.36722v1 Announce Type: new Abstract: Large language model (LLM) agents repeatedly load reusable content, such as skills, documents, and memory entries, into the current context. Re-encoding this content for every request wastes computation. Position-independent caching (PIC) alleviates this by encoding each artifact independently and reusing its key-value (KV) states at arbitrary positions, but it incurs a quality loss relative to full-context prefill. Existing methods repair this loss by restoring global position IDs or recomputing selected tokens. In this work, we isolate the source of the loss, finding that the positional mismatch has minor effect, and independently cached artifacts retain faithful representations: reading a provided artifact stays largely accurate, and perfo
קרא במקור המקורי