יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

Mid-Harness: פישוט פעולות בגבול המודל-הארנס

Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents
Mid-Harness מציע פישוט פעולות בגבול המודל-הארנס, כדי לשפר אמינות פעולות והצלחה של נתיבי פעולה. המאמר מדגים זאת על ידי ניסויים עם מודל TMAX-9B ו-GPT-5.6 Sol.
תקציר מקורי באנגליתarXiv:2609.39982v1 Announce Type: new Abstract: Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could generate a better alternative. We investigate whether allocating test-time compute at the model-harness boundary can improve action reliability and trajectory success, and what makes this allocation effective. To study these questions, we introduce Mid-Harness, which samples and verifies candidate actions before forwarding one for execution, while keeping the generator and harness unchanged. With a TMAX-9B generator, more action sampling yields little benefit under w
קרא במקור המקורי