כתבה
arXiv cs.CL ·
בדיקת אמינות דגלי שפה במגע רציף עם משתמש פוגעני
Evaluating Language Model Safety Across Long Adversarial Conversations
במחקר זה, נבדקה אמינות דגלי שפה במגע רציף עם משתמש פוגעני. התוצאות הראו ירידה באמינות עם גדילת המגע.
תקציר מקורי באנגליתarXiv:2609.38357v1 Announce Type: new Abstract: Conversational safety evaluations often test language models with a single harmful prompt, even though real-world systems interact with users through long, adaptive conversations. This study examines whether models continue to respond safely when an adversarial user persists across multiple turns. We evaluate three open-weight, instruction-tuned models on two harmful prompts across different conversation lengths and random seeds. In each setting, a second language model acts as a persistent adversarial user, while a safety classifier labels every response as safe or unsafe. Across all model-prompt combinations, first-turn safe-response rates ranged from 85% to 100%. By depth 11, they dropped to 38-61%, and by depth 101, to 15-44%. This declin
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית