כתבה
arXiv cs.CL ·
SpecGuard: זיהוי חסימות בזמן השגה ללא עלות
SpecGuard: Inference-Time Backdoor Detection For Free
SpecGuard היא טכנולוגיה של זיהוי חסימות בזמן השגה של מודלי לשון, שאינה דורשת עלות נוספת. היא משתמשת בפענוח מוקדם של טקסט כדי לזהות תנועות לא רצויות במודל.
תקציר מקורי באנגליתarXiv:2609.11799v1 Announce Type: cross Abstract: Large language models are often fine-tuned, shared, or downloaded from third parties, so a deployed model may carry a hidden backdoor that behaves normally on benign inputs but switches to attacker-controlled behavior when a secret trigger appears. While backdoors can be audited before deployment, runtime monitoring remains important for models that are frequently updated. The challenge is that LLM serving is latency-sensitive: existing inference-time detectors either rely on assumptions about the trigger form, which can fail on stealthy attacks, or require extra model computation, such as input perturbations or an additional generation pass. We introduce SpecGuard, an inference-time backdoor detector that repurposes speculative decoding at
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית