יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מתאימים לחוקה: התערבויות בזמן השגה למניעת פגיעה ושימוש לא ראוי

Constitutional adapters: Inference-time interventions for misalignment and misuse
טכניקה חדשה לאימון מודלי AI לקיים עקרונות קבועים, ולהגן על עצמה מפני שימוש לא ראוי. הטכניקה, שנקראת 'מתאימים לחוקה', מציעה דרך נוספת להגן על AI מפני פגיעה ושימוש לא ראוי.
תקציר מקורי באנגליתarXiv:2609.36657v1 Announce Type: cross Abstract: Training models to act in accordance with an explicitly defined set of principles, or "constitution," has shown promise as a robust and transparent mechanism for AI alignment. However, the generality and flexibility of such methods remain unclear. Here, we show that constitution-consistent behavior can be distilled from synthetic corpora into lightweight objects (low-rank adapters and steering vectors). Despite never seeing a harmful request or jailbreak during training, such objects increase jailbreak defense success and measured alignment -- particularly at long context lengths and against multi-turn attacks, where they outperform both prompted and steered baselines. Subtracting control-trained from constitution-trained objects further ac
קרא במקור המקורי