יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

LTBD: גבולות אמון למניעת התקפות זריקת פרומפט

LTBD: Learnable Trust-Boundary Delimiters for Prompt Injection Defense
LTBD הוא חידוש להגנה על מודלי שפה גדולים (LLM) מפני התקפות זריקת פרומפט. הוא משתמש בגבולות אמון כדי להבדיל בין הוראות משתמש מהימנות לנתונים חיצוניים לא מהימנים. LTBD מוכיח עצמו כיעיל נגד התקפות מותאמות.
תקציר מקורי באנגליתarXiv:2610.11634v1 Announce Type: cross Abstract: Large language models (LLMs) perform remarkably well on complex tasks, yet remain highly vulnerable to prompt injection attacks, where malicious instructions embedded in external data can override user intent. Existing defenses remain limited by model fine-tuning requirements, vulnerability to adaptive attacks, or reliance on brittle handcrafted prompts. We argue that a fundamental source of this vulnerability is the lack of an explicit representation of trust provenance. To address this, we introduce Learnable Trust-Boundary Delimiters (LTBD), a lightweight defense that explicitly encodes trust boundaries in the input while keeping the LLM parameters unchanged. LTBD uses a small number of learnable delimiters to distinguish trusted user in
קרא במקור המקורי