יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

HarnessBandit: תוכנית למידה-העברה משותפת לקידום רכיבי רגשיות

HarnessBandit: Joint Learnability-Transferability Scheduling for Multi-Harness Agentic Reinforcement Learning
HarnessBandit היא תוכנית למידה-העברה משותפת שמטרתה לשפר את רכיבי הרגשיות של רשתות עצמאיות. היא משתמשת בשיטת GRPO ובסקירת גרדיאנטים נמוכי-ממדים. התוכנית נבחנה על ידי Qwen3.5-2B והראתה תוצאות טובות יותר מאשר קידום מרובה-רשתות.
תקציר מקורי באנגליתarXiv:2609.13739v1 Announce Type: cross Abstract: Language-model agents are increasingly deployed through diverse harnesses that differ in system prompts, tool schemas, control loops, and trajectory formats. The same model can perform unevenly across these interfaces, making robustness to harness variation an important objective. A natural approach is to train a shared policy through multiple harnesses, but doing so introduces a scheduling problem: each training step should favor a harness that currently provides a useful learning signal while also producing an update that benefits the other harnesses. We develop HarnessBandit, an online scheduler that selects one harness per optimizer step. After a group-relative policy optimization (GRPO) update, it observes learnability -- the mean abso
קרא במקור המקורי