יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

Meta-SecAlign: טירונות LLMs נגד התקפות פרומפט-אינג'קציה לשם רובוסטיות

Meta-SecAlign: Training LLMs against Prompt Injection for Robust Agents
טירונות LLMs נגד התקפות פרומפט-אינג'קציה לשם רובוסטיות. Meta-SecAlign מציע תרגילי אימון ובדיקות חדשים לשם טירונות LLMs נגד התקפות פרומפט-אינג'קציה.
תקציר מקורי באנגליתarXiv:2507.02735v4 Announce Type: replace-cross Abstract: Prompt injection attacks, where untrusted data contains an injected prompt to manipulate the system, have been listed as the top security threat to AI agents. By fine-tuning on simulated prompt injections, SecAlign, a leading open defense, reports LLMs with good test-time robustness and negligible benign utility drop. By scaling up training and evaluations, however, we find that SecAlign actually suffers from significant utility degradation, especially in agentic tasks where the threat of prompt injection is prominent. Motivated by this, we propose Meta-SecAlign for utility-preserving defense by (1) randomized injection position during training to avoid shortcut learning and (2) self-generated responses as high-quality in-distributi
קרא במקור המקורי