יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

RaReCache: גשר להפחתת זמן טעינה במערכות LLM

RaReCache: Bridging the Gap in Cross-Model KV Cache Reuse via Rank disagreement-based Selective Recomputation
RaReCache היא פלטפורמה שמסייעה למערכות LLM לצמצם את זמן הטעינה. היא משתמשת בשיטת רקומפוטציה מבוזרת כדי לאפשר לדגמים קטנים לטעון את הקונטקסט של דגמים גדולים. RaReCache יכולה לצמצם את זמן הטעינה בכ-3.04 פעמים.
תקציר מקורי באנגליתarXiv:2610.11358v1 Announce Type: cross Abstract: Cross-model KV-cache reuse remains a key challenge in modern LLM serving. Coding agents and multi-model systems increasingly route a shared context across models: a user may switch models mid-session, or a cascade may escalate a difficult query. Because KV caches contain model-specific representations, each switch typically forces the receiving model to prefill the entire context from scratch. Recent work shows that closed-form linear maps can translate KV caches between models in the same family, but transfer accuracy degrades as the model-size gap widens. In this paper, we establish that these transfer failures are concentrated in a small subset of information-dense tokens. To bridge this gap, we introduce RaReCache, a framework that enab
קרא במקור המקורי