כתבה
arXiv cs.LG ·
TwinRouterBench: בדיקת רוטרים סטטיים ודינמיים לרוטרים LLM ריאליסטיים
TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing
מוצג TwinRouterBench, בדיקת רוטרים חדשה לרוטרים LLM ריאליסטיים, המספקת בדיקה סטטית ודינמית. הבדיקה כוללת 970 פרפקסים של רוטרים ובודקת את היכולת שלהם לבחור במודלים זולים יותר שישמרו את ההצלחה של המשימה. הבדיקה נעשית באמצעות תוכנה שנכתבה באנגלית ומספקת תוצאות ניסויים ובדיקה.
תקציר מקורי באנגליתarXiv:2605.18859v4 Announce Type: replace Abstract: LLM routing matters most in long-horizon applications such as coding agents, deep research systems, and computer-use agents, where a single user request triggers many model calls. Routing each call to the cheapest sufficient model can cut costs without sacrificing quality, yet existing router benchmarks evaluate routers only on one-shot prompts. They never expose the router-visible prefix at an intermediate agent step, never test whether a cheaper replacement preserves downstream task success, and often rely on online LLM judges at evaluation time. We introduce TwinRouterBench, a step-level routing benchmark with two tracks. The static track provides 970 router-visible prefixes from 520 instances across SWE-bench, BFCL, mtRAG, QMSum, and
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית