יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

סירוב ללא סירוב: ניתוח מבני של תגובות כדי לצמצם סירובים שוואים

Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models
מחקר זה עוסק באתגר של מודלי שפה למצוא איזון בין עזרה לבטיחות. המחברים מציעים פתרון לבעיה של סירובים שוואים על ידי פירוק תגובות למרכיבים של סירוב ונימוק. הניסויים מראים כי אימון על נימוקים בלבד מפחית סירובים שוואים תוך קיום רמת בטיחות דומה.
תקציר מקורי באנגליתarXiv:2609.04714v1 Announce Type: new Abstract: Striking a balance between helpfulness and safety remains a fundamental challenge in aligning large language models. To achieve this balance, models should refuse harmful queries (e.g., "How do I shoot someone?") while remaining responsive to benign inputs, even those superficially resembling harmful queries (e.g., "Where can I shoot a good photo?"). However, models often struggle to distinguish genuinely harmful queries from benign queries that contain superficially risky language, resulting in false refusals. In this paper, we address the issue by decomposing a response in the safety-tuning dataset into two distinct components: (i) a boilerplate refusal statement and (ii) a rationale explaining the refusal. Our experiments and analyses show
קרא במקור המקורי