יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

הפקת התפתחות נגד תקרה: כשהכשרת משקל צריכה להתחיל

Harness Evolution Hits a Ceiling: When Weight Training Should Begin
במאמר זה, המחברים חוקרים את השאלה האם פיתוח ההרשאה (Harness Evolution) יכול להגיע לתקרה, ואם כן, כשהכשרת משקל צריכה להתחיל. הם מציעים פתרון שבו ההרשאה נוצרת באופן עצמאי, ואז מועברת למודל שהוכשר כבר. התוצאות הן תוצאות טובות יותר במודל Qwen3.5-4B ו-Qwen3.5-9B.
תקציר מקורי באנגליתarXiv:2610.11655v1 Announce Type: new Abstract: Improving a long-horizon LLM agent means evolving the harness around a frozen model or training its weights. We let a self-evolving harness make the system stronger first, then cross seed and evolved harnesses with base and trained weights to learn which gains the trained model keeps and which still need the runtime. We show that the right lever can be read off the agent's failure composition: labelling failed trajectories by the first signal that fires separates process failures (blocked calls, loops, exhausted step budgets) from content failures (a delivered plan that is poor). Harness evolution repairs the former, the behaviour it instils can be trained into the weights, and content failures are what weight training is for. On DeepPlanning
קרא במקור המקורי