כתבה
arXiv cs.AI ·
MARBO: אופטימיזציה רלציונלית של תכונות רציונליות לאג'נטים LLM במשחקי סדוקציה חברתית
MARBO: Relational Belief Grounding for LLM Agents in Social Deduction Games
MARBO היא תכונת אופטימיזציה של LLM אג'נטים במשחקי סדוקציה חברתית, המבוססת על תכונות רלציונליות. היא מאפשרת לאג'נטים ללמוד להגיב באופן רציונלי למצבים חברתיים.
תקציר מקורי באנגליתarXiv:2609.06563v1 Announce Type: new Abstract: Social deduction games (SDGs) require agents to reason under partial observability by maintaining relational beliefs about hidden roles and team alignments. While recent LLM-agent approaches improve gameplay through prompting and preference optimization, they often optimize actions and in-game speech without explicitly grounding them in such beliefs. This frequently leads to strategically inconsistent behavior, especially for compact LLM agents. We introduce Multi-Agent Relational Belief Optimization (MARBO), a belief-grounded preference optimization framework that leverages relational beliefs to guide strategic decisions and in-game speech. MARBO provides preference feedback only when behaviors are supported by reliable relational beliefs an
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית