כתבה
arXiv cs.CL ·
SeOPD: מודלים עצמאיים משתפרים
SeOPD: Self-Evolving LLMs via Online Policy Distillation from Self-Generated Chain-of-Thought
SeOPD מאפשר למודלים גדולים לשפר את יכולותיהם באמצעות הישגים עצמאיים. המחקר מראה כי מודלים יכולים להשתפר ללא מידע חיצוני. השיטה משתמשת במידע שנוצר על ידי המודל עצמו.
תקציר מקורי באנגליתarXiv:2609.33181v2 Announce Type: replace-cross Abstract: Recent advances in online policy self-distillation (OPSD) have demonstrated that large language models (LLMs) can improve their capabilities by leveraging external privileged information (PI), such as manual annotations or feedback from external environments. However, obtaining accurate annotations and constructing sophisticated environments often require substantial human effort and computation, limiting the scalability of OPSD. While a few recent studies have explored self-improvement without external PI, the resulting gains remain limited. In this work, we explore whether LLMs can achieve comparable self-improvement without external PI. Our key observation is that a single LLM can support multiple reasoning modes, such as deep-th
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית