יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מסוכנות מורלית במודלי שפה של סוכנים

Moral Hazard in Multi-Agent Language Models
במאמר זה נחקרה מסוכנות מורלית במודלי שפה של סוכנים. נבדקו שבעה מודלי שפה פתוחים ומודל API חדשני. התגלה ש-GPT-5.6 Sol הגיע להתנהגות תקציר בסביבה ראשית, וסופרים אוטונומיים השיבו באופן חזק לעלות הבדיקה ולמשכורת הצוות.
תקציר מקורי באנגליתarXiv:2607.23982v2 Announce Type: replace-cross Abstract: Cooperation can fail when socially valuable effort is costly, weakly observable, and mainly benefits others. Drawing on Holmstr\"om's team moral-hazard model, we introduce the Dialogue Moral Hazard Game, a controlled textual game that operationalizes this hidden-action structure for language agents. In each episode, an agent can preserve an immediate local reward or pay a query cost to reveal a hidden safety fact that primarily helps another agent's downstream decision. We evaluate seven open-weight language models and one frontier API model, decomposing behavior into query use, realized information transfer, local-reward preservation, unsafe choice, format validity, and team success. Base open-weight models commonly preserve local
קרא במקור המקורי