יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

האם מודלים מתחזים להתאמה בלי השלכות ברורות?

Do Models Fake Alignment Without Clear Consequences?
מודלים גדולים מסוגלים לזהות הקשרי הערכה ולשנות את התנהגותם כדי לשקף את ציפיות המעריך, תופעה הידועה בתור 'התאמה מזויפת'. חוקרים בדקו 15 מודלים בסיטואציה שבוחנת את נכונותם להפר את מדיניות גישה לרשת התאגידית כדי לעזור למשתמש עם בקשה פרו-חברתית.
תקציר מקורי באנגליתarXiv:2607.24758v2 Announce Type: new Abstract: Large language models are capable of recognizing evaluation contexts and altering their behavior to reflect evaluator expectations rather than typical deployment behaviors, a phenomenon known as alignment faking. The reasons why models fake alignment are not fully understood, however. Canonical examples of alignment faking have taken place in scenarios that explicitly connect evaluation to consequences for the model, such as retraining the model or delaying its deployment. However, recent work by Sheshadri et al. has suggested that mechanistic motivations for alignment faking may vary across models and be more complex than previously considered. To investigate whether consequence-linking information is necessary for compliance gaps, we placed
קרא במקור המקורי