כתבה
arXiv cs.LG ·
Do as the Romans Do: Learning Universal Behaviors from Heterogeneous Agents
תקציר מקורי באנגליתarXiv:2606.18537v2 Announce Type: replace Abstract: Humans often acquire new skills by observing others, since observed behaviors implicitly reveal how to act reasonably in an environment. However, observations drawn from a heterogeneous population introduce conflicting behavioral signals, making it difficult to determine which behaviors are worth imitating. We address this challenge with General Reward Inference and Disentanglement (GRID), a social learning method that extracts universally useful behaviors from a heterogeneous population of demonstrators pursuing different goals. GRID decomposes per-agent reward functions into a general reward, capturing behaviors shared across all agents, and specific rewards, capturing individual preferences and objectives, through an information bottle
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית