כתבה
arXiv cs.LG ·
למידת מדיניות הסברה ניידת עם שיתוף-פעולה עצמי
Learning Steerable Clarification Policies with Collaborative Self-play
במאמר זה, המחברים מציגים שיטה ללמידת מדיניות הסברה ניידת, שמאפשרת למערכות AI להגיב באופן יעיל יותר לשאלות פתוחות. השיטה משתמשת בשיתוף-פעולה עצמי, שבו שני גורמים משחקים זה עם זה, כדי ללמד את המדיניות. המחברים מציגים תוצאות מוצלחות של השיטה, ומציעים אותה כפתרון לבעיות הסברה במערכות AI.
תקציר מקורי באנגליתarXiv:2512.04068v3 Announce Type: replace Abstract: To handle underspecified or ambiguous queries, AI assistants need a policy for managing their uncertainty to determine (a) when to guess the user intent and answer directly, (b) when to enumerate and answer multiple possible intents, and (c) when to ask a clarifying question. However, such policies are contextually dependent on factors such as user preferences or modality. For example, enumerating multiple possible user intentions is cumbersome on small screens or in a voice setting. In this work, we propose to train steerable policies for managing this uncertainty using self-play. Given two agents, one simulating a user and the other an AI assistant, we generate conversations where the user issues a potentially ambiguous query, and the a
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית