יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

Rufus-Air: מתכון אימון פוסט-LLM פתוח

Rufus-Air: An Open LLM Post-Training Recipe
Rufus-Air הוא מתכון אימון פוסט-LLM פתוח וניתן לשחזור, המבוסס על מודל GLM-4.5-Air-Base. המתכון כולל 8 שלבים, החל מ-SFT וכלה ב-RLHF. התוצאות מראות ש-Rufus-Air משפר על הגרסה הרשמית של GLM-4.5-Air.
תקציר מקורי באנגליתarXiv:2609.29421v2 Announce Type: replace-cross Abstract: Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and RLHF. We document the data, reward design, infrastructure, stage order, and stagewise results needed to reproduce the recipe. Stages progress from basic to advanced capabilities and from hard, verifiable rewards to softer judge-based signals. Training builds on open-source components and public data, much of it used as released, without new human annotation or an in-house distillation teacher. Our main findings are that (i) diverse, high-quality SFT establishes a strong capability floor; (ii) d
קרא במקור המקורי