יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

Scepsy: תפעול זרמי עבודה של סוכנים באמצעות פיפליינים של LLM

Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines
Scepsy היא מערכת שמספקת זרמי עבודה של סוכנים באמצעות פיפליינים של LLM. היא מספקת תפעול גבוה וזמן תגובה נמוך יותר בהשוואה למערכות אחרות.
תקציר מקורי באנגליתarXiv:2604.15186v3 Announce Type: replace-cross Abstract: Agentic workflows carry out complex tasks by orchestrating multiple large language models (LLMs) and tools. Serving them at a target throughput with low latency is hard because they are written in arbitrary agentic frameworks and their execution times are unpredictable: execution branches, fans out, or recurs in data-dependent ways. Since their LLMs often outnumber the available GPUs, they also oversubscribe GPUs. We describe Scepsy, a serving system that schedules arbitrary multi-LLM agentic workflows onto a GPU cluster. Scepsy exploits the insight that, while the end-to-end latency of an agentic workflow is unpredictable, each LLM's fraction of execution time is comparatively stable across requests. Scepsy profiles each LLM under
קרא במקור המקורי