יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

המאבק בין המשך וסירוב: ניתוח מכניסטי של הפריצה המונעת-המשך ב-LLMs

The Struggle Between Continuation and Refusal: A Mechanistic Analysis of the Continuation-Triggered Jailbreak in LLMs
במאמר זה, חוקרים חשפים את הסיבות העמוקות לחולשות הבטיחות של LLMs. הם מציעים רעיון להגברת הבטיחות בזמן הריצה.
תקציר מקורי באנגליתarXiv:2603.08234v2 Announce Type: replace-cross Abstract: With the rapid advancement of large language models (LLMs), the safety of LLMs has become a critical concern. Despite significant efforts in safety alignment, current LLMs remain vulnerable to jailbreaking attacks. However, the root causes of such vulnerabilities are still poorly understood, necessitating a rigorous investigation into jailbreak mechanisms across both academic and industrial communities. In this work, we focus on a continuation-triggered jailbreak phenomenon, whereby simply relocating a continuation-triggered instruction suffix can substantially increase jailbreak success rates. To uncover the intrinsic mechanisms of this phenomenon, we conduct a comprehensive mechanistic interpretability analysis at the level of att
קרא במקור המקורי