יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

עבר מונג' אל תבניות בטיחות

Beyond Refusal Patterns: Safe-Role Internalization for Robust and Generalizable LLM Safety Alignment
SSRFT הוא גישה חדשה להתאמת בטיחות למודלי שפה גדולים. היא מתמקדת בהנחלת תפקיד בטוח ולא בסירוב. הניסויים הראו תוצאות טובות יותר מגישות קודמות.
תקציר מקורי באנגליתarXiv:2610.07023v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved remarkable capabilities but remain vulnerable to jailbreak attacks that elicit harmful or unsafe outputs. Existing safety alignment approaches, including Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), often require substantial attack-specific supervision and computational resources, while remaining susceptible to shallow safety alignment and over-refusal. To address these challenges, we introduce SSRFT(Supervised Safe-Role Fine-Tuning), the first framework that reformulates safety alignment as the internalization of a predefined safe role. SSRFT constructs a Safe-Role Question-Answer (SRQA) dataset from psychometric questions, limited jailbreak prompts, and a sa
קרא במקור המקורי