יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

הגנות על מודלים LLM פתוחים רגישות לתקיפות פשוטות

Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks
חוקרים מצאו כי הגנות על מודלים LLM פתוחים רגישות לתקיפות פשוטות. התקיפות, שאינן מבוססות על אופטימיזציה, יכולות לעקוף הגנות אלו ולאפשר שימוש תוקפני במודלים. המחקר מציע פתרון לבעיה זו.
תקציר מקורי באנגליתarXiv:2605.26526v2 Announce Type: replace Abstract: Recent defenses for safeguarding open-weight large language models (LLMs) are intended to prevent adversarial usage. Underlying these defenses is an assumption that new harmful behavior is learned through fine-tuning rather than elicited by jailbreaking the model. Yet, pretrained LLMs already encode substantial harmful knowledge across many domains, which raises an important question: can an adversary jailbreak safeguarded models, to achieve harmful usage without fine-tuning at all? In this paper, we show that open-weight safeguards are susceptible to simpler strategies that, despite being well known, have not been systematically evaluated against these safeguards. Specifically, we evaluate two low-cost attacks--abliteration and prefillin
קרא במקור המקורי