יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

FinCUABuild: יכולתם של סוכנים לבנות מדדים יציבים לשימוש חישובי כספים דינאמיים?

FinCUABuild: Can Agents Build Reliable Benchmarks for Dynamic Financial Computer Use?
סוכנים יכולים לבנות מדדים יציבים לשימוש חישובי כספים דינאמיים? המחקר חוקר את היכולת של סוכנים לבנות תרגילי בדיקה יעילים לשימוש חישובי כספים. המחברים מציגים את FinCUABuildBench, מדד לבדיקת יכולת הבנייה של סוכנים, ואת FinCUABuildAgent, מערכת של סוכנים לבניית תרגילי בדיקה דינאמיים. התוצאות מציגות שסוכנים יכולים לבנות תרגילי בדיקה יעילים לשימוש חישובי כספים.
תקציר מקורי באנגליתarXiv:2609.07603v1 Announce Type: new Abstract: Financial scenarios are diverse and complex, spanning varying data conditions, tool configurations, and workflows. Yet existing CUA, Computer-Using Agent, evaluation tasks remain largely manually constructed, limiting scalable coverage of real-world financial scenarios. Then, can agents autonomously construct diverse CUA evaluation tasks for financial scenarios? Evaluating this capability poses three key challenges: scenario coverage of construction requests, fair comparison across construction methods, and reliable assessment of generated task quality. To solve these, we introduce FinCUABuildBench, a benchmark for evaluating financial CUA task construction, featuring: (i) 576 construction requests covering 24 financial workflows and three ty
קרא במקור המקורי