כתבה
arXiv cs.AI ·
RSI-Master: סטרוקטורה של ניסויים להנחיה של שיפור עצמי אוטונומי של דגמים
RSI-Master: Structuring Experiments to Guide Autonomous Model Improvement
RSI-Master מאפשר שיפור עצמי אוטונומי של דגמים דרך ניסויים מאורגנים. הוא עובד עם דגמי Qwen ו-GPT-5, ומשפר את היכולות שלהם. RSI-Master גם עובד עם דגמי Instruct, ומשפר את היכולות שלהם. הוא גם עובד עם דגמי HorizonMath, ומשפר את היכולות שלהם. RSI-Master יכול לשפר את היכולות של דגמים עצמאית, וללא תלות באדם.
תקציר מקורי באנגליתarXiv:2609.35561v2 Announce Type: replace Abstract: Recursive self-improvement (RSI) seeks to enable AI systems to participate in improving their own capabilities. A concrete pathway is autonomous model development, where agents iteratively explore post-training strategies to improve a base model. This setting faces two challenges: agents may exploit open-ended experimental actions through hacking, and repeated experimentation may lead to strategy lock-in, where an early direction is refined rather than reconsidered. We introduce RSI-Master, which addresses the two challenges at two levels: regularize step-wise actions, avoiding hacking behaviors, and promote well-structured exploration of research directions, avoiding strategy lock-in. RSI-Master consists of an Experiment OS, which enable
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית