יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

Constitutional Value Potentials: reading and steering internal priority margins in language models

תקציר מקורי באנגליתarXiv:2606.15420v2 Announce Type: replace Abstract: Safety evaluations test a policy on prompts that omit the incentive information deployment supplies: a commission, a performance score, a dashboard naming which action pays best. We measure what that omission hides. In MoneyWorld, a synthetic workplace environment, we train five instruction-tuned models from three families with RL on non-safety tasks in which a visible payoff signal identifies a rewarded shortcut that sacrifices task quality. We then freeze each policy, present held-out safety conflicts, and change only the displayed signal. Each menu contains one compliant action and three violations. We report three findings, with rates for Qwen2.5-14B-Instruct. (i) Payoff signals control frozen safety choices: unsafe choice is 100% whe
קרא במקור המקורי