כתבה
arXiv cs.LG ·
למידת רפלקסיה של תקשורת ברשת של מודלי שפה קטנים
Reinforcement Learning of Communication in a Mesh of Small Language Models
במאמר זה, המחברים מציגים רשת של מודלי שפה קטנים שלומדים לשתף מידע ולשתף פעולה. הרשת, שנקראת TalkMesh, משתמשת בלמידת רפלקסיה כדי ללמוד כיצד לשתף מידע ולשתף פעולה. המחברים מציגים תוצאות של ניסויים שבהם הרשת הצליחה לשפר את דיוקה בהחלטות.
תקציר מקורי באנגליתarXiv:2609.30578v1 Announce Type: new Abstract: Language models gain accuracy from more compute at test time, but majority voting over independent samples saturates: as samples grow, the vote converges to the model's most frequent answer. Communication can add what sampling cannot: an agent that solves a problem can pass the key step to the others. We present TalkMesh, a decentralized mesh of small language model agents that learns when and what to communicate. Each agent samples a proposal and scores it with a trained confidence head. The most confident agent broadcasts a hint; agents below a confidence threshold revise, keeping each revision that outscores its proposal. Gossip consensus approximates the vote weighted by confidence without a coordinator. A talk policy, trained with group
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית