יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

AlcaTRAz - תרגיל עץ-שלטון נגד פריצות תאגיד

AlcaTRAz - Anchored Tree-Rule Defense Against Jailbreaks
מאמר חדש מציג תרגיל חדש נגד פריצות תאגיד במודלי שפה גדולים. התרגיל, AlcaTRAz, פועל רק על הטקסט המובא ולא דורש שינויים או אימון מחדש של המודל. AlcaTRAz יוצר תרגיל תעבורה נייד שמכנס פרטורבציות תווים ברמת התו, ומפריע לסדריות רגולריות שהפריצות ניצולות. המאמר מדגים את התרגיל על 33 מודלי שפה פתוחים, 22 סוגי פריצות ובנק אופן של שאלות קצרות.
תקציר מקורי באנגליתarXiv:2609.03693v1 Announce Type: cross Abstract: Large language models (LLMs) are vulnerable to jailbreak attacks that bypass safety alignment through carefully crafted prompts. Many existing defenses require access to model weights or internals, making them difficult to apply to black-box deployments. We propose AlcaTRAz (Anchored Tree-Rule defense Against jailbreaks), a prompt-level defense based on rule trees that operates exclusively on the input text and requires no modification or retraining of the target model. The method automatically learns a transferable transformation rule that inserts controlled character-level perturbations at selected positions, thereby disrupting structural regularities exploited by jailbreak attacks while largely preserving the model's utility on benign qu
קרא במקור המקורי