יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

אופטימיזציה של דינמיקות הטריינינג של GRPO למודלי שפה קטנים

GRPO Training Dynamics for Small Language Models
במאמר זה, נחקרה אופטימיזציה של דינמיקות הטריינינג של GRPO למודלי שפה קטנים. המחקר כולל חקירה של השפעת גודל הקבוצה על סיכויי ההתכווצן של המדיניות, יציבות הטריינינג וביצועי המדיניות.
תקציר מקורי באנגליתarXiv:2609.39321v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has emerged as a memory-efficient reinforcement fine-tuning (RFT) technique for reasoning-intensive tasks. How- ever, GRPO training dynamics on small language models (SLMs) remain poorly understood, limiting its reliable adoption and reproducibility in open and resource- constrained environments. In this work, we present a systematic study of GRPO fine-tuning for SLMs ranging from 1.5B to 7B parameters under a practical single- node 8xA100 compute budget. Our study spans multiple model families and reasoning domains, including mathematics, coding, and multiple-choice question answering (MCQ) in science. Across these settings, we analyze how group size affects policy convergence, training stability, an
קרא במקור המקורי