יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

הגנה על LLMs באמצעות סימני בטיחות סמויים

Safeguarding LLMs via Model-Agnostic Latent Safety Signals from Dark Knowledge
לפיתוח סימני בטיחות סמויים עבור LLMs, כדי להגן עליהם מפני פגיעה. החידוש נעשה על ידי השוואת תפוקות של תבניות פגיעה ותבניות בטיחותיות. המאמר עוסק בפיתוח כלי לזיהוי והגנה על LLMs. הכלי, הנקרא LADE, עושה שימוש בסימני בטיחות סמויים שנמצאו בתפוקות של LLMs. LADE יכול להגן על LLMs מפני פגיעה, כולל פגיעה על ידי תבניות פגיעה.
תקציר מקורי באנגליתarXiv:2610.07532v1 Announce Type: cross Abstract: LLMs have advanced rapidly, raising growing concerns about their safety. Recent work has proposed approaches to detect and defend against attacks including defenses at decoding stage that leverage models' hidden states. However, existing decoding-stage defenses suffer from two limitations. First, they introduce a trade-off between safety and over-refusal, where strengthening safety degrades the model's helpfulness on benign queries. Second, many of these methods rely on internal hidden states and are thus restricted to specific architectures, incurring substantial overhead and limited generalization across models. To address these limitations, we introduce LADE (Latent Safety Signals for Defense), which leverages latent safety signals extra
קרא במקור המקורי