כתבה
arXiv cs.AI ·
SEATauBench: Progressively Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages
תקציר מקורי באנגליתarXiv:2606.28715v2 Announce Type: replace-cross Abstract: While AI development and evaluation for Southeast Asia (SEA) has grown rapidly, agent capabilities in regional languages are still poorly understood despite its importance to sovereign AI. To fill this gap, we introduce SEATauBench, the first agent-focused evaluation framework for SEA sovereign AI. It adapts Tau2-Bench to five languages---Mandarin, Vietnamese, Thai, Indonesian, and Filipino---and evaluates agents across progressively localized settings that vary the language of user-agent interaction, tool specifications, and task domains. Across three models, we find that English agent capabilities transfer reasonably well when only the conversation language changes, but quality and robustness degrade sharply as more task contexts
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית