יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

SparKV: טווח-מודע KV Cache Loading לביצועי LLM Efficient On-Device

SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference
SparKV הוא פרקטיקה של טעינת KV שמשפרת את ביצועי ה-LLM במכשירים ניידים. SparKV משלב זרימת KV מהענן עם חישוב על המכשיר. הוא מודל את עלויות ה-KV הפרטיות ומחליט האם לטעון או לחשב כל חתיך KV, תוך כדי חיתוך של שני הדרכים להפחתת זמן תגובה.
תקציר מקורי באנגליתarXiv:2604.21231v3 Announce Type: replace-cross Abstract: Efficient inference for on-device Large Language Models (LLMs) remains challenging due to limited hardware resources and the high cost of the prefill stage, which processes the full input context to construct Key-Value (KV) caches. We present SparKV, an adaptive KV loading framework that combines cloud-based KV streaming with on-device computation. SparKV models the cost of individual KV chunks and decides whether each chunk should be streamed or computed locally, while overlapping the two execution paths to reduce latency. To handle fluctuations in wireless connectivity and edge resource availability, SparKV further refines offline-generated schedules at runtime to rebalance communication and computation costs. Experiments across d
קרא במקור המקורי