כתבה
arXiv cs.AI ·
SynCo: סינתזה של נתונים ל-LLM
SynCo: Data Synthesis Co-Training for Self-Evolving LLMs via Multi-Agent Reinforcement Learning
SynCo הוא כלי לסינתזה של נתונים ל-LLM, המשתמש בלמידת חיזוק רב-סוכנית. הוא מאפשר ל-LLM להתפתח באופן עצמאי, תוך כדי למידה מניסיון. SynCo רואה עלייה משמעותית בביצועים לעומת שיטות אחרות.
תקציר מקורי באנגליתarXiv:2610.11345v1 Announce Type: new Abstract: Self-evolving LLM agents promise to improve autonomously through continual interaction and learning, reducing their dependence on manually curated supervision. Realizing this promise requires not only updating the agent, but also evolving its training experience as its capabilities change. However, most existing pipelines rely on static datasets or separately updated synthesis models, causing previously useful tasks to become trivial while overly difficult tasks remain uninformative. This growing mismatch between agent capability and training experience limits sustained self-improvement. To address this problem, we propose SynCo, an agentic data synthesis co-training framework for self-evolving LLMs based on multi-agent reinforcement learning
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית