יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

Miles v0.1: מערכת לאימון פוסט-ייצור

Miles v0.1: Production-Level Post-Training
Miles v0.1 היא מערכת מלאה לאימון פוסט-ייצור, הבנויה על עקרונות של ניקיון ואישור. המערכת תומכת באימון RL, LoRA RL, ואימון עם דגמי דיפוזיה. Miles פותחה על ידי radixark ומוצעת כקוד פתוח.
תקציר מקורי באנגליתarXiv:2609.08368v1 Announce Type: new Abstract: We present Miles v0.1, a full-stack, production-ready system for frontier post-training. Building upon the clean design of slime, Miles designs each stage of the reinforcement-learning (RL) training loop around a single principle: components should be verified, clean, and customizable. With accuracy, efficiency, reliability, and scalability as first-class goals, Miles aims to make frontier-scale RL accessible to researchers and enterprises alike. This report walks through the system end to end: rollout engines built on SGLang, a trainer with a choice of two backends (NVIDIA Megatron-LM and PyTorch FSDP), and three weight-synchronization transports for different deployment topologies. Beyond full-parameter RL, Miles also supports LoRA RL, on-p
קרא במקור המקורי