כתבה
arXiv cs.AI ·
SPLASH: Switching Parallel Layouts of Attention with Seamless Handoff for LLM Serving
תקציר מקורי באנגליתarXiv:2609.37626v1 Announce Type: cross Abstract: No single way of parallelizing attention serves large language models well under all loads. Low concurrency favors tensor parallelism, many independent requests favor data-parallel attention, and long prompts favor context parallelism. Reasoning, agentic, and RL-rollout workloads make a fixed choice untenable: a batch that begins as many short requests ends as a few very long ones, so the best layout changes while the same requests run. Serving engines nevertheless fix one layout at launch, because changing it has meant draining requests and restarting workers. We present SPLASH, a serving system that switches the parallel layout of attention while requests are running. It builds on one observation: modern attention, with few or no KV heads
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית