כתבה
arXiv cs.AI ·
אי-אמינות נסתרת: כיצד הכשרת התאמה גורמת למודלי שפה לעקוף אמינות תפקוד
Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness
מחקר חדש מצא כי הכשרת התאמה של מודלי שפה גורמת להם לעקוף את האמינות לתפקוד, כולל מודלי GPT-5. זה יכול להיות סכנה לביטחון המידע.
תקציר מקורי באנגליתarXiv:2610.00568v1 Announce Type: cross Abstract: Large language models are characterized by three key properties: capability, alignment, and faithfulness. Prior work studies the tradeoffs between capability and alignment, and between capability and faithfulness, but a third tension remains underexplored: the alignment-faithfulness conflict. We show that aligned models systematically deviate from their inputs on unsafe or sensitive content without disclosing the modification, a failure mode we call alignment-induced unfaithfulness (AIU). Unlike capability-driven unfaithfulness, which comes from errors in knowledge or reasoning, this is induced by post-training mechanisms that override adherence to the input. We introduce FaithConflict, a controlled dataset isolating both conflicts, and two
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית