כתבה
arXiv cs.AI ·
Character Training for Risk-Averse Agents
תקציר מקורי באנגליתarXiv:2609.38093v1 Announce Type: new Abstract: Risk aversion in resources could prevent misaligned AI agents from causing catastrophic harm. Misaligned but risk-averse agents would tend to favor safer strategies like making deals with humans over riskier strategies like rebelling. We train agents to be risk averse through character training, finding that persona traits provide a robust mechanism for instilling risk preferences. To do this, we construct a model constitution describing constant absolute risk aversion (CARA) over an agent's resources and instill it through on-policy distillation. Despite never seeing the benchmark's decision format during training, character-trained models are competitive with baselines trained directly on it, and generalise better than them out of distribut
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית