כתבה
arXiv cs.AI ·
EVA-Bench: פלטפורמה חדשה לבדיקת זיכרון קולי
EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents
EVA-Bench היא פלטפורמה חדשה לבדיקת זיכרון קולי של זיכרון קולי. היא כוללת 213 תצורות בשלושה תחומי עסקים, ובודקת את יכולתן של 12 מערכות שונות.
תקציר מקורי באנגליתarXiv:2605.13841v3 Announce Type: replace-cross Abstract: Voice agents are increasingly deployed across enterprise applications. However, no existing benchmark jointly addresses realistic conversation simulation and comprehensive voice-specific evaluation. We present EVA-Bench, an end-to-end evaluation framework that addresses both. On the simulation side, EVA-Bench orchestrates dynamic bot-to-bot audio conversations with automatic simulation validation that detects user simulator error and appropriately regenerates conversations before scoring. On the measurement side, EVA-Bench introduces two composite metrics: EVA-A (Accuracy) and EVA-X (Experience). EVA-Bench includes 213 scenarios across three enterprise domains, a controlled perturbation suite for accent and noise robustness, and mul
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית