כתבה
arXiv cs.AI ·
לסרב על לא לסרב: ניתוח סטרוקטורלי של תגובות טונינג-בטוחות להפחתת סירובים שקריים בדגמי שפה
Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models
במאמר זה, המחברים מציעים שיטה חדשה להפחתת סירובים שקריים בדגמי שפה על ידי הכשרת רציונלים.
תקציר מקורי באנגליתarXiv:2609.04714v1 Announce Type: cross Abstract: Striking a balance between helpfulness and safety remains a fundamental challenge in aligning large language models. To achieve this balance, models should refuse harmful queries (e.g., "How do I shoot someone?") while remaining responsive to benign inputs, even those superficially resembling harmful queries (e.g., "Where can I shoot a good photo?"). However, models often struggle to distinguish genuinely harmful queries from benign queries that contain superficially risky language, resulting in false refusals. In this paper, we address the issue by decomposing a response in the safety-tuning dataset into two distinct components: (i) a boilerplate refusal statement and (ii) a rationale explaining the refusal. Our experiments and analyses sh
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית