כתבה
arXiv cs.AI ·
SRD-GUARD: תשתית הגנה למודלי LLM דרך תיקון סמנטי והערכת סיכון משותפת לחשיפת כוונה נסתרת
SRD-GUARD: A Defense Framework of LLMs via Semantic Rewriting and Joint Multi-Model Scoring for Latent Intent Exposure
תשתית הגנה למודלי LLM דרך תיקון סמנטי והערכת סיכון משותפת. SRD-GUARD מציע פתרון להגנה על LLMs מפני תקיפות כלאבק. התשתית נבחנה בעבודה על Llama-3-8B-Uncensored ו-DeepSeek-V4-Flash.
תקציר מקורי באנגליתarXiv:2609.06540v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in safety-critical applications, yet jailbreak attacks can conceal harmful intent through role-playing, fictional scenarios, or seemingly benign motivations. Existing inference-time defenses may miss disguised attacks or excessively refuse legitimate requests. We propose SRD-GUARD, a parameter-free, black-box defense framework that exposes concealed intent through semantic rewriting and consensus-based risk assessment. Given an input prompt, SRD-GUARD generates five semantically related rewrites that preserve the underlying objective while removing unnecessary contextual packaging. The original prompt and rewrites are jointly evaluated by multiple independent LLM-based safety scorers on
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית