כתבה
arXiv cs.LG ·
MARBO: אופטימיזציה של אמונות יחסיות ל-LLM Agents
MARBO: Relational Belief Grounding for LLM Agents in Social Deduction Games
MARBO היא שיטה לאופטימיזציה של אמונות יחסיות עבור LLM Agents במשחקים חברתיים. היא מאפשרת לסוכנים לקבל החלטות אסטרטגיות על בסיס אמונות יחסיות, ולשפר את ביצועיהם במשחקים אלה. הקוד של MARBO זמין ב-GitHub.
תקציר מקורי באנגליתarXiv:2609.06563v1 Announce Type: cross Abstract: Social deduction games (SDGs) require agents to reason under partial observability by maintaining relational beliefs about hidden roles and team alignments. While recent LLM-agent approaches improve gameplay through prompting and preference optimization, they often optimize actions and in-game speech without explicitly grounding them in such beliefs. This frequently leads to strategically inconsistent behavior, especially for compact LLM agents. We introduce Multi-Agent Relational Belief Optimization (MARBO), a belief-grounded preference optimization framework that leverages relational beliefs to guide strategic decisions and in-game speech. MARBO provides preference feedback only when behaviors are supported by reliable relational beliefs
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית