כתבה
arXiv cs.LG ·
ActiveSaddler: אופטימיזציה אוטומטית של קורס ולמידה להשפעת סוכן
ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization
ActiveSaddler מציגה אופטימיזציה אוטומטית של קורס ולמידה להשפעת סוכן, כדי לשפר את יכולות LLM. השיטה משתמשת בבנדיט לא-תחבירי כדי לקבוע את הסצנריות המשמעותיות ביותר לאופטימיזציה. ActiveSaddler נבחנה על GAIA2 ו-Terminal-Bench 2.0 והראתה תוצאות טובות יותר מאשר אופטימיזציה קבועה.
תקציר מקורי באנגליתarXiv:2610.00906v1 Announce Type: cross Abstract: Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback. However, existing methods primarily optimize how the harness is updated while largely fixing which training scenarios generate the feedback that drives those updates. As the harness evolves, the scenarios most useful for further optimization can change, suggesting that the training curriculum itself should adapt alongside the harness. We formulate this missing dimension of harness optimization as an automated curriculum learning problem and introduce ActiveSaddler. ActiveSaddler models the evolving curriculum as a non-stationary bandit with dynamically instantiated optimizatio
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית