יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

עבר מונג' אל תפקידים בטוחים ל-LLM

Beyond Refusal Patterns: Safe-Role Internalization for Robust and Generalizable LLM Safety Alignment
SSRFT הוא כלי חדש להבטחת בטיחות LLM. הוא מאפשר תפקידים בטוחים ומגן על המודלים מפני התקפות jailbreak. ה-SSRFT נבדק על מודלים שונים והראה תוצאות טובות.
תקציר מקורי באנגליתarXiv:2610.07023v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable capabilities but remain vulnerable to jailbreak attacks that elicit harmful or unsafe outputs. Existing safety alignment approaches, including Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), often require substantial attack-specific supervision and computational resources, while remaining susceptible to shallow safety alignment and over-refusal. To address these challenges, we introduce SSRFT(Supervised Safe-Role Fine-Tuning), the first framework that reformulates safety alignment as the internalization of a predefined safe role. SSRFT constructs a Safe-Role Question-Answer (SRQA) dataset from psychometric questions, limited jailbreak prompts, and a safe
קרא במקור המקורי