כתבה
arXiv cs.AI ·
Mid-Harness: הגברת אמינות פעולות בין מודל להרנס
Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents
Mid-Harness משפר את אמינות הפעולות בין מודל להרנס. הוא מנסה ומאמת פעולות לפני ביצוע, תוך שמירה על המודל וההרנס ללא שינוי. עם TMAX-9B, Mid-Harness מגביר את הצלחת הפעולות.
תקציר מקורי באנגליתarXiv:2609.39982v1 Announce Type: cross Abstract: Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could generate a better alternative. We investigate whether allocating test-time compute at the model-harness boundary can improve action reliability and trajectory success, and what makes this allocation effective. To study these questions, we introduce Mid-Harness, which samples and verifies candidate actions before forwarding one for execution, while keeping the generator and harness unchanged. With a TMAX-9B generator, more action sampling yields little benefit under
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית