כתבה
arXiv cs.LG ·
Measuring Reward-Seeking via Contrastive Belief Updates
תקציר מקורי באנגליתarXiv:2607.18966v1 Announce Type: cross Abstract: Language models trained with reinforcement learning may learn to optimize the grader's judgment rather than the intended objective. This "reward-seeking" is difficult to measure because a model that pursues the grader's judgment and one that pursues the intended objective behave identically whenever the grader rewards the intended behavior. We measure reward-seeking using Contrastive Synthetic Document Finetuning to change a model's beliefs about what the grader rewards, putting those beliefs in conflict with what users or developers want, and measuring the rate at which the model adopts each party's preferred behavior. Applied to intermediate checkpoints of a capabilities-focused OpenAI o3 RL run, without safety training, we find that thes
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית