יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

Suan: תיקון בטיחות ישיר של נטיות חברתיות בדגמי שפה גדולים

Suan: Rectifying Direct Preference Safety Alignment in Large Language Models
Suan הוא אלגוריתם חדשני לאופטימיזציה של נטיות חברתיות לדגמי שפה גדולים. הוא נועד לשפר את הבטיחות של הדגמים ולמנוע תגובות שליליות. Suan פותח על ידי צוות של מדעני דטא ומהנדסי תוכנה.
תקציר מקורי באנגליתarXiv:2609.08634v2 Announce Type: replace-cross Abstract: Integrating robust safety guardrails into Large Language Models (LLMs) is essential for delivering helpful yet harmless responses. While proprietary systems exhibit reliable safety controls, their underlying methodologies and trade-offs remain largely undisclosed. Achieving comparable security in open-weight models remains a persistent challenge, as post-trained variants frequently suffer from over-refusal and degraded general quality. To overcome these drawbacks, we introduce Suan, a novel preference optimization algorithm. Unlike existing methods, we formulate the optimization objective directly at the gradient level, bypassing the standard variational derivation. As a result, we obtain more interpretable and robust training dynam
קרא במקור המקורי