כתבה
arXiv cs.LG ·
ביצועים, יעילות וקריסה - יתרונות ואתגרים
Performance, Efficiency and Collapse -- Advantages and Challenges in Offline Post-training of Code LLMs
מחקר זה בודק אם אימון עם למידת חיזוק (RL) יכול להתבצע באופן מלא אופליין, תוך שימוש במאגרי נתונים קיימים. התוצאות מראות שניתן לשפר משמעותית את ביצועי ייצור קוד של מודלים גדולים, תוך שימוש בשעות ספורות של אימון.
תקציר מקורי באנגליתarXiv:2609.11956v1 Announce Type: new Abstract: Post-training with reinforcement learning (RL) is a critical phase in the development of code-generating large language models (LLMs), as it ensures adherence to instructions and the production of functionally correct code. This process typically requires computationally intensive code sample generation from Transformer-based LLMs and substantial GPU-CPU communication for sequence verification. To address these computational challenges, this work examines whether RL-based post-training can be performed entirely offline by leveraging existing datasets rather than generating new samples. The findings indicate that, with only a few hours of training, zero-shot code generation performance of LLMs can be substantially improved without online sampl
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית