כתבה
arXiv cs.LG ·
TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing
תקציר מקורי באנגליתarXiv:2605.18859v3 Announce Type: replace Abstract: LLM routing matters most in long-horizon applications such as coding agents, deep research systems, and computer-use agents, where a single user request triggers many model calls. Routing each call to the cheapest sufficient model can cut costs without sacrificing quality, yet existing router benchmarks evaluate routers only on one-shot prompts. They never expose the router-visible prefix at an intermediate agent step, never test whether a cheaper replacement preserves downstream task success, and often rely on online LLM judges at evaluation time. We introduce TwinRouterBench, a step-level routing benchmark with two tracks. The static track provides 970 router-visible prefixes from 520 instances across SWE-bench, BFCL, mtRAG, QMSum, and
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית