כתבה
arXiv cs.AI ·
UserProxyBench: בדיקת מדדי נאמנות של סימולטורי משתמש לבניית סטנדרטים ולימוד
UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training
UserProxyBench היא בדיקה של מדדי נאמנות של סימולטורי משתמש לבניית סטנדרטים ולימוד. המאמר עוסק בבדיקת סימולטורי משתמש של LLM עבור בניית סטנדרטים ולימוד. הבדיקה כוללת בדיקה של סימולטורי משתמש של GPT-5.
תקציר מקורי באנגליתarXiv:2609.38043v1 Announce Type: new Abstract: Interactive agent benchmarks and multi-turn reinforcement learning increasingly place a second language model in the role of the user. This simulated user controls what information the agent receives and when, yet current benchmarks score only the agent and do not directly measure whether the user correctly executed its assigned role. We introduce UserProxyBench, an evaluation layer over the tau-bench family, and the User Fidelity Score (UFS), which measures adherence to the benchmark's private user instructions using task-grounded rubric criteria scored independently of agent success. Holding the agent fixed at GPT-5.5 and varying only the user proxy across 375 enterprise tasks changes mean task reward by 15.2 points, while 24.4% of successf
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית