יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

סיכון מוסרי במודלים שפה רב-סוכנים

Moral Hazard in Multi-Agent Language Models
חוקרים בדקו את התנהגותם של מודלים שפה רב-סוכנים במצבים של סיכון מוסרי. הם השתמשו במודלים כגון OLMo-7B ובשיטות כגון fine-tuning ו-GEPA prompt optimization. התוצאות הראו שהמודלים יכולים לשפר את התנהגותם, אך לא תמיד באופן שמשפר את השיתוף פעולה.
תקציר מקורי באנגליתarXiv:2607.23982v1 Announce Type: cross Abstract: Cooperation can fail when socially valuable effort is costly, weakly observable, and mainly benefits others. Drawing on Holmstr\"om's team moral-hazard model, we introduce the Dialogue Moral Hazard Game, a controlled textual game that operationalizes this hidden-action structure for language agents. In each episode, an agent can preserve an immediate local reward or pay a query cost to reveal a hidden safety fact that primarily helps another agent's downstream decision. We evaluate seven open-weight language models and decompose behavior into query use, realized information transfer, local-reward preservation, unsafe choice, format validity, and team success. Base models commonly preserve local reward without team success or query without c
קרא במקור המקורי