יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

TwinRouterBench: רוטר סטטי ודינמי מהיר לבדיקה של רוטינג LLM ריאליסטי

TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing
TwinRouterBench הוא תקן לבדיקת רוטינג LLM. הוא כולל שני עקבות: סטטי ודינמי. העקבות הסטטי מספק 970 ראשי תיבות של רוטר שניתן לראות, כולל 520 ראשי תיבות של רוטר שניתן לראות מ-5 סטים שונים. העקבות הדינמי מספק תשתית שמרצה רוטרים על 500-קס SWE-bench Verified. TwinRouterBench תומך באיטרציות אפליין מהירות ובדיקה תחת רצוף רוטיני.
תקציר מקורי באנגליתarXiv:2605.18859v3 Announce Type: replace-cross Abstract: LLM routing matters most in long-horizon applications such as coding agents, deep research systems, and computer-use agents, where a single user request triggers many model calls. Routing each call to the cheapest sufficient model can cut costs without sacrificing quality, yet existing router benchmarks evaluate routers only on one-shot prompts. They never expose the router-visible prefix at an intermediate agent step, never test whether a cheaper replacement preserves downstream task success, and often rely on online LLM judges at evaluation time. We introduce TwinRouterBench, a step-level routing benchmark with two tracks. The static track provides 970 router-visible prefixes from 520 instances across SWE-bench, BFCL, mtRAG, QMSum
קרא במקור המקורי