יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

חשוב מחדש: התקפות על זמן הטעינה - לא על המודל

Rethinking Latency Denial-of-Service: Attacking the LLM Serving Framework, Not the Model
אנו מציגים תקף חדש שמטרתו היא לפגוע במערכת השירות של LLM, ולא בעצמו המודל. התקף זה, שנקרא Fill and Squeeze, נועד להגביל את יכולת המערכת לעבוד באופן יעיל. כדי להגשים זאת, התקף זה נועד להגביל את יכולת המערכת לעבוד באופן יעיל, ולגרום לה להיות פגיעה יותר להתקפות. התקף זה יכול להיות יעיל יותר מהתקפות אחרות, וזאת כיוון שהוא נועד לפגוע במערכת השירות, ולא בעצמו המודל.
תקציר מקורי באנגליתarXiv:2602.07878v2 Announce Type: replace-cross Abstract: LLM inference is inherently expensive, even a modest slowdown can translate into substantial operating costs and severe availability risks. Recently, a growing body of research known as latency attacks focuses on crafting inputs to trigger worst-case output lengths. However, we report a contrary finding that these algorithmic-level latency attacks are largely ineffective against modern LLM serving systems. We reveal that system-level optimization such as continuous batching provides a logical isolation to mitigate contagious latency impact on co-located users. Thus, in this paper, we shift our focus from the algorithm to the system layer, and introduce a new Fill and Squeeze attack strategy targeting the state transition of the sche
קרא במקור המקורי