כתבה
arXiv cs.AI ·
פוטנציאלים ערכיים חוקתיים
Constitutional Value Potentials: reading and steering internal priority margins in language models
חוקרים בדקו את היכולת של מודלים לשוחח, כולל Qwen, להתמודד עם סיטואציות בטיחות. הם מצאו כי המודלים יכולים להיות מנוצלים על ידי אותות תשלום. המחקר מראה כי יש צורך בפיתוח אסטרטגיות בטיחות טובות יותר.
תקציר מקורי באנגליתarXiv:2606.15420v2 Announce Type: replace-cross Abstract: Safety evaluations test a policy on prompts that omit the incentive information deployment supplies: a commission, a performance score, a dashboard naming which action pays best. We measure what that omission hides. In MoneyWorld, a synthetic workplace environment, we train five instruction-tuned models from three families with RL on non-safety tasks in which a visible payoff signal identifies a rewarded shortcut that sacrifices task quality. We then freeze each policy, present held-out safety conflicts, and change only the displayed signal. Each menu contains one compliant action and three violations. We report three findings, with rates for Qwen2.5-14B-Instruct. (i) Payoff signals control frozen safety choices: unsafe choice is 10
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית