כתבה
arXiv cs.LG ·
ביטחון שכבת ההפניה: הגנה מפני פעולות תקיפה והשימוש הפסול במשאבי התשתית
Inference-Layer Security: Defending Against Adversarial Inference and Infrastructure Abuse
במאמר זה נחקר הביטחון של שכבת ההפניה והגנה מפני פעולות תקיפה והשימוש הפסול במשאבי התשתית. המחברים פיתחו דטסט ומודל לזיהוי פעולות תקיפה.
תקציר מקורי באנגליתarXiv:2609.38239v1 Announce Type: cross Abstract: A Technical Report: Operating a large language model (LLM) as a service requires more than inference infrastructure: the provider must also defend against adversarial interactions that seek to exploit the service, including jailbreaking for harmful use, sophisticated denial of service, and distillation attacks. We study this problem at the inference layer, using a hypothetical frontier lab, Five Elements Inc., as a running example. Because no public labelled dataset of adversarial LLM usage exists, we introduce a structural causal model (SCM) that generates a realistically grounded, labelled dataset of user-sessions, with coordinated multi-account campaigns, platform feedback, and three tiers of label observability. On this dataset we train
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית