כתבה
arXiv cs.LG ·
מדידת הפער באפשרות למשימות קטנות: מתי סלמ רגיל יכול להיות מספיק להארנס של סוכן?
Measuring the Microtask Eligibility Gap: When Is an Off-the-Shelf SLM Enough for an Agent Harness?
במאמר זה נבחן את הפער בין סלמ קטנים לאפשרות לבצע משימות קטנות. נמצא כי סלמ רגילים לא עומדים בתנאים לבצע משימות אלו. נבחן גם את השפעת קיצור זיכרון על התוצאות.
תקציר מקורי באנגליתarXiv:2610.00025v1 Announce Type: cross Abstract: Agent harnesses increasingly want to run small language models (SLMs) on the microtasks around a frontier large language model (LLM) planner: auto-approving shell commands, writing memory, selecting tools, ranking past turns. We ask whether off-the-shelf SLMs meet practitioner-defined thresholds and, when they fail, why, and whether quantization changes the answer. We build a benchmark of 4 such microtasks with fixed prompts and automatic metrics, each with a pre-specified threshold $\tau$ anchored to a cheap non-LLM baseline and a CI-aware eligibility rule (a configuration passes only if its confidence bound clears $\tau$). Sweeping Qwen3 0.6/1.7/4/8B at their best (FP16, greedy, one frozen prompt, no tuning), we find an eligibility gap: 0
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית