כתבה
arXiv cs.CL ·
RL-ADA: פלטפורמה להכשרת סוכני דיון עם תגובה-עולם להתמודדות עם תקיפה
RL-ADA: A World-Feedback Framework for Adversarially Robust Enterprise Dialogue Agents
פלטפורמה חדשה להכשרת סוכני דיון עם תגובה-עולם, שמספקת תגובה-עולם לסוכני דיון ומאפשרת הכשרה עם תגובה-עולם. הפלטפורמה נועדה לסוכני דיון שמטרתם לספק תגובות עם תגובה-עולם, ולא רק לספק תגובות. הפלטפורמה כוללת שלושה חלקים: חלק ראשון, שבו נבחן הסוכן הדיוני; חלק שני, שבו נבחן הסוכן הדיוני עם תגובה-עולם; וחלק שלישי, שבו נבחן הסוכן הדיוני עם תגובה-עולם ועם תגובה-עולם.
תקציר מקורי באנגליתarXiv:2609.02902v1 Announce Type: new Abstract: Deploying task-oriented dialogue agents in enterprise customer support faces a persistent annotation bottleneck: robust training requires labelled interaction data at scale, yet enterprise conversational logs are privacy-sensitive and expensive to annotate, while user behaviour evolves faster than labelling pipelines can keep pace. We present RL-ADA (Reinforcement Learning with Adversarial Dialogue Agents), a co-evolutionary training framework that eliminates this bottleneck by replacing human labels with \emph{world feedback}: consequence-based reward signals derived directly from measurable interaction outcomes. A Customer Support Agent (DA, 3B parameters) and an Adversarial Customer Agent (CA, 7B parameters) co-evolve in an adversarial are
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית