כתבה
arXiv cs.AI ·
המאבק בין המשך והסרבה: ניתוח מניפולטיבי של הפריצה המוצא-המשך ב-LLMs
The Struggle Between Continuation and Refusal: A Mechanistic Analysis of the Continuation-Triggered Jailbreak in LLMs
במאמר זה, חוקרים חוקרים את הפריצה המוצא-המשך ב-LLMs ומציעים פתרון לבעיית הבטיחות ב-LLMs. המאמר כולל ניתוח מניפולטיבי של הפריצה והצעה לפתרון.
תקציר מקורי באנגליתarXiv:2603.08234v2 Announce Type: replace Abstract: With the rapid advancement of large language models (LLMs), the safety of LLMs has become a critical concern. Despite significant efforts in safety alignment, current LLMs remain vulnerable to jailbreaking attacks. However, the root causes of such vulnerabilities are still poorly understood, necessitating a rigorous investigation into jailbreak mechanisms across both academic and industrial communities. In this work, we focus on a continuation-triggered jailbreak phenomenon, whereby simply relocating a continuation-triggered instruction suffix can substantially increase jailbreak success rates. To uncover the intrinsic mechanisms of this phenomenon, we conduct a comprehensive mechanistic interpretability analysis at the level of attention
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית