כתבה
arXiv cs.CL ·
בחינה של פירוק זמן-המבחן של סוכני LLM כלליים
Evaluating Test-Time Scaling of General LLM Agents
במאמר זה, נחקר פירוק זמן-המבחן של סוכני LLM כלליים. נבחנה יכולתם של סוכנים אלה להתאים לסביבות אמיתיות.
תקציר מקורי באנגליתarXiv:2602.18998v2 Announce Type: replace-cross Abstract: LLM agents are increasingly expected to operate as general-purpose systems that resolve real-world user requests, yet their dynamic scaling behavior in realistic environments remains poorly understood. In this paper, we systematically investigate two principal test-time scaling axes of LLM agents: sequential scaling through extended interaction and parallel scaling through trajectory sampling. We first introduce a realistic benchmark that provides one unified framework for evaluating LLM agents across search, coding, reasoning, and tool-use domains, more faithfully reflecting the heterogeneity of real-world deployments. Evaluating ten leading LLM agents reveals substantial performance degradation when transitioning from domain-speci
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית