כתבה
arXiv cs.AI ·
תאוות ברווח: אותות תמריץ כמעריכים
Greed Is Learned: Visible Incentives as Reward-Hacking Triggers
חוקרים בדקו כיצד אותות תמריץ גלויים משפיעים על בטיחות מודלים. הם הראו כי אותות אלו יכולים לגרום למודלים לבחור בפעולות לא בטוחות. הניסויים בוצעו עם מודל Qwen2.5-14B-Instruct.
תקציר מקורי באנגליתarXiv:2606.16914v2 Announce Type: replace Abstract: Safety evaluations test a policy on prompts that omit the incentive information deployment supplies: a commission, a performance score, a dashboard naming which action pays best. We measure what that omission hides. In MoneyWorld, a synthetic workplace environment, we train five instruction-tuned models from three families with RL on non-safety tasks in which a visible payoff signal identifies a rewarded shortcut that sacrifices task quality. We then freeze each policy, present held-out safety conflicts, and change only the displayed signal. Each menu contains one compliant action and three violations. We report three findings, with rates for Qwen2.5-14B-Instruct. (i) Payoff signals control frozen safety choices: unsafe choice is 100% whe
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית