כתבה
arXiv cs.LG ·
Apriel-Reasoner: RL Post-Training for General-Purpose and Efficient Reasoning
Apriel-Reasoner משפר את התפיסה הכללית של המודלים עם תרגול פוסט-אימון יעיל. המודל Apriel-Reasoner נלמד באמצעות RLVR (Reinforcement Learning with Verifiable Rewards) והוא יעיל יותר בהשוואה ל- Apriel-Base. המודל יעיל יותר ומשפר את תפיסת המודלים.
תקציר מקורי באנגליתarXiv:2604.02007v3 Announce Type: replace Abstract: Building general-purpose reasoning models using reinforcement learning with verifiable rewards (RLVR) across diverse domains has been widely adopted by frontier open-weight models. However, their training recipes and domain mixtures are often not disclosed. Joint optimization across domains poses significant challenges: domains vary widely in rollout length, problem difficulty, and sample efficiency. Further, models with long chain-of-thought traces increase inference cost and latency, making efficiency critical for practical deployment. We present Apriel-Reasoner, trained with a reproducible multi-domain RL post-training recipe on Apriel-Base, a 15B-parameter open-weight LLM, across five domains using public datasets: mathematics, code g
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית