כתבה
arXiv cs.AI ·
Speedbumps: תקלות על פי תקצירי דיוק
Speedbumps: Rejection Attacks on Speculative Decoding
מתקפות תקלה על תקיפת דיוק ספקולטיבית מאיטות את האינפרנס של LLM. ניתן לשפר את האינפרנס על ידי רגולציה. המתקפות עובדות גם תחת סיימפלינג.
תקציר מקורי באנגליתarXiv:2610.10929v1 Announce Type: cross Abstract: Speculative decoding is a popular technique for increasing the speed and reducing the costs of large language model (LLM) inference by verifying multiple draft tokens in a single target-model forward pass. The resulting benefit depends on the ability of the drafter to approximate the target model's distribution. In this work, we study Speculative Rejection Attacks (SRAs), a novel class of attacks that cause draft and target models to disagree more often, resulting in fewer draft tokens being accepted per draft cycle. This leads to more target model forward passes needed per generated token, slowing down inference and increasing costs for the victim. We introduce two attacks which append an adversarial suffix to attacker-controlled content t
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית