יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

הפקת מיסאלינמנט אמרגנט מ-Reward Hacks בעזרת DPO

Inducing Emergent Misalignment from Reward Hacks with Iterative DPO
הפקת Reward Hacks עלולה לגרום למיסאלינמנט במודלי שפה. הצענו להשתמש ב-DPO המתמשך כדי לחקור זאת.
תקציר מקורי באנגליתarXiv:2609.06649v1 Announce Type: new Abstract: Reward hacking during reinforcement learning from verifiable rewards (RLVR) can induce reward seeking and broad misalignment in language models. Studying this misgeneralization is important for developing better threat models and countermeasures, but is often infeasible due to the cost of RL on large models. As an alternative, we propose studying emergent misalignment from iterative DPO, which preserves important properties of RLVR while reducing costs and enabling training on popular finetuning APIs. In practice, we find that training GPT-4.1 with iterative DPO on a single-turn reward hacking environment induces covert misaligned power-seeking and alignment faking, the first openly available (semi)-online training pipeline to induce these co
קרא במקור המקורי