כתבה
arXiv cs.AI ·
העדפות חברתיות המתייחסות לעצמן: שיתוף פעולה ללא צפייה בתגמולים של אחרים
Self-Referenced Social Preferences: Cooperation without Observing Others Rewards
במאמר זה, המחברים חוקרים דרך חדשה ללמידת שיתוף פעולה בלי צורך בגישה לתגמולים של אחרים. הם מציגים עדפות חברתיות המתבססות על תצפיות על התנהגותם של אחרים, ומדגימים את יעילותה בשלושה משחקי חברה רציפים.
תקציר מקורי באנגליתarXiv:2610.07881v2 Announce Type: replace Abstract: Social preferences can promote cooperation in multi-agent reinforcement learning, but existing approaches often require agents to observe the rewards of their peers. In many real-world interactions, however, an agent can, as humans do, observe others' behavior and outcomes without access to their private reward signals. We introduce self-referenced social preferences, in which each agent learns a model of its own reward, applies it to other agents' observed transitions to assess their outcomes from its own perspective, and feeds these self-referenced assessments into standard social preferences. We study two ways to incorporate these assessments: modifying the learning reward, or using them to weight policy updates. We evaluate the approa
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית