יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

דגמים קטנים: חקרי טבע טבעיים לשונות מדדית ב-GRPO

Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
דגמים קטנים יותר מציגים שונות מדדית גבוהה יותר, המובילה לדיוק טוב יותר במבחני סיבוכיות מתמטית. כולל שימוש ב-GPT-5.
תקציר מקורי באנגליתarXiv:2605.30789v3 Announce Type: replace Abstract: We identify a new dimension for enhancing rollout diversity in Group Relative Policy Optimization (GRPO) for LLMs. While GRPO relies on diverse rollouts, prevailing strategies primarily increase diversity by injecting more token-level randomness, which may introduce step-wise noise and lead to incoherent trajectories. We uncover that smaller models within the same model family inherently exhibit higher policy-level diversity, indicated by their superior pass@k relative to larger counterparts as sample counts increase. Unlike token-level noise, this diversity is temporally correlated, preserves logical consistency, and provides structured exploration signals for gradient estimation. We thus propose S2L-PO (Small-to-Large Policy Optimizatio
קרא במקור המקורי