יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

AgBench: בנצ'מרקים למערכות AI אוטונומיות

AgBench: Agentic AI Benchmarks for Personal AI Devices
AgBench הוא סוויטת בנצ'מרקים למערכות AI אוטונומיות על מכשירים אישיים. הוא בודק ביצועים מקומיים, היברידיים וענניים. התוצאות מראות שמכשירים אישיים יכולים לבצע משימות רבות מקומית, אך הביצועים נמוכים יותר מאשר ביצועים ענניים.
תקציר מקורי באנגליתarXiv:2609.38652v1 Announce Type: new Abstract: Agentic AI systems increasingly rely on cloud-hosted large language models for planning, tool use, and iterative execution, raising concerns about API cost and data exposure. Advances in personal AI devices enable agents to execute locally, but limited resources on device may affect task success and performance. Existing benchmarks are inadequate for systematically characterizing these trade-offs across devices, workloads, and deployment architectures. We present AgBench, a benchmark suite and open artifacts for reproducible evaluation of agentic AI on personal devices. Using AgBench, we evaluate local, hybrid, and cloud execution across agentic workloads, examining task success, latency, cloud API cost, and data exposure. Our results, drawn
קרא במקור המקורי