יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

למה פריצות צלחו בדגמי שפה דיפוזיים: ניתוח נוף אנרגטי

Why Jailbreaks Succeed in Diffusion Language Models: An Energy Landscape Analysis
פריצות לדגמי שפה דיפוזיים: ניתוח נוף אנרגטי. חוקרים חקרו את סיבות הצלחתן של פריצות בדגמי שפה דיפוזיים. הם גילו שפריצות עובדות על ידי חציית גבול אנרגטי. הם פיתחו שלושה סימני דטקציה שאינם תלויים בהכשרה. הסימנים חקרו את דגמי LLaDA-8B, LLaDA-1.5, Dream-7B ו-LLaDA-MoE-7B.
תקציר מקורי באנגליתarXiv:2609.30841v1 Announce Type: new Abstract: Existing attacks and defenses for diffusion-based large language models (dLLMs) target specific vulnerabilities but lack a shared framework explaining why attacks succeed. We propose one by interpreting safety alignment as shaping the denoising energy landscape: a well-aligned model routes harmful queries toward safe outputs through an energy barrier that separates the two regions. Current jailbreak attacks reduce to two strategies for circumventing this barrier: obscuring the query's safety disposition at initialisation, or intervening mid-trajectory to force the denoising path across the energy barrier. From this perspective and the result that masked diffusion models minimise kinetic energy during denoising, we derive three complementary,
קרא במקור המקורי