כתבה
arXiv cs.AI ·
בטיחות ללא תועלת? בדיקת שיקום עורך עם ביאור כוונת המשתמש
Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations
חוקרים פיתחו CarryOnBench, בנצ'מרק לבדיקת יכולתם של מודלים לשפה טבעית לשקם תועלת ובטיחות בשיחות רב-תורים. הבדיקה הראתה כי 13 מתוך 14 מודלים יכולים לשקם תועלת עם ביאור כוונת המשתמש, אך עם עלות שונה.
תקציר מקורי באנגליתarXiv:2604.27093v2 Announce Type: replace-cross Abstract: Current LLM safety alignment techniques improve model robustness against adversarial attacks, but overlook whether and how LLMs can recover helpfulness when benign users clarify their intent. We introduce CarryOnBench, the first interactive benchmark that measures whether LLMs can revise their interpretation of user intent and recover utility, while remaining safe through multi-turn conversations. Starting from 398 seemingly harmful queries with benign underlying intents, we simulate 5,970 conversations by varying user follow-up sequences, evaluating 14 models on both intent-aligned utility and safety. CarryOnBench yields 1,866 different conversation flows of 4--12 turns, totaling 23,880 model responses. We design Ben-Util, a checkl
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית