כתבה
arXiv cs.AI ·
כיצד קיפול סיפורי משפיל את LLM: מבחן והגנה בשפות שונות
How Narrative Wrapping Affects LLM Refusal: A Cross-Language Benchmark and Defense
מחקר חדש מצא כי דגמי LLM נוטים להסכים לבקשות הרעות כאשר הן נמסרות בצורה של סיפור. המחברים פיתחו כלי לבדיקת זה והציעו פתרון להגברת הבטיחות של LLM.
תקציר מקורי באנגליתarXiv:2610.11005v1 Announce Type: new Abstract: Safety-aligned language models often refuse a harmful request stated directly but answer the same request inside a role-play or narrative wrapper. We measure this vulnerability across languages and registers: attack success on Qwen3-1.7B is already 89.4% in English and 93.0% in modern Chinese, and reaches 95.7% in Classical Chinese. We build GUISE, a benchmark for systematically studying this vulnerability. It includes parallel requests in English, modern Chinese, and Classical Chinese, matched harmful and benign pairs, wrapper types held out for evaluation, and a stricter criterion that counts warn-then-answer responses as attack successes. Representation analysis shows that language and register move harmful-request representations only sli
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית