כתבה
arXiv cs.CL ·
Optimizing Sparse Outcomes Through Dense Behavioral Signals via Value-Guided Preference Distillation
תקציר מקורי באנגליתarXiv:2609.14648v1 Announce Type: new Abstract: Aligning multi-turn dialogue agents is usually framed as matching turn-level human preferences, yet direct optimization of long-term outcomes is often ineffective and prone to reward hacking. We formulate long-horizon dialogue optimization as a multi-objective reinforcement learning problem and train a multi-head value model that predicts a vector of observed user behaviors across multiple look-ahead horizons. Our findings demonstrate that a scalarized composite of dense auxiliary behavioral signals enables effective credit assignment and optimization of sparse outcomes. However, optimizing unconstrained single-objective proxies might induce policy degradations that are harmful when the agent is exposed to real users. To identify these failur
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית