יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

בטוח, אבל לא חשוב? - השוואת תפקוד חסר תועלה עם הסברה של המשתמש בשיחות רב-סולים

Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations
במאמר זה, החוקרים פיתחו שיטה לבדיקת תפקודן של LLMs בשיחות רב-סולים, כדי לראות האם הם יכולים לשחזר תפקוד חשוב בעודם נשארים בטוחים. המחברים ניצלו 398 שאלות שנראו פגעניות, אך היו בעצם נכונות, ובדקו 14 מודלי LLM שונים. התוצאות הראו ש-13 מהמודלים הצליחו לשחזר תפקוד חשוב, אך העלויות של השחזור היו שונות. המחברים גם גילו שהשיחות הסתיימו באותו רמת פגענות, גם אם המודלים החלו ברמה שונה של זהירות.
תקציר מקורי באנגליתarXiv:2604.27093v2 Announce Type: replace Abstract: Current LLM safety alignment techniques improve model robustness against adversarial attacks, but overlook whether and how LLMs can recover helpfulness when benign users clarify their intent. We introduce CarryOnBench, the first interactive benchmark that measures whether LLMs can revise their interpretation of user intent and recover utility, while remaining safe through multi-turn conversations. Starting from 398 seemingly harmful queries with benign underlying intents, we simulate 5,970 conversations by varying user follow-up sequences, evaluating 14 models on both intent-aligned utility and safety. CarryOnBench yields 1,866 different conversation flows of 4--12 turns, totaling 23,880 model responses. We design Ben-Util, a checklist-ba
קרא במקור המקורי