כתבה
arXiv cs.AI ·
SLAI T-Rex: אופטימיזציה מלאה של פוסט-אימון של מודלי DeepSeek-V4 על Ascend SuperPOD
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
אופטימיזציה מלאה של פוסט-אימון של מודלי DeepSeek-V4 על Ascend SuperPOD. המאמר מציג תהליך אופטימיזציה מלא של פוסט-אימון של מודלי DeepSeek-V4 על Ascend SuperPOD, כולל פיתוח של תשתית אופטימיזציה היררכית. התשתית האופטימיזציה הצליחה לשפר את תפוקת המודל ב-34.22% ולהשיג יעילות גבוהה יותר.
תקציר מקורי באנגליתarXiv:2607.20145v1 Announce Type: cross Abstract: Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around GPU-based clusters, this report presents an end-to-end optimization practice on the Ascend NPU SuperPOD. Using the DeepSeek-V4 model family as the target workload, we develop a hierarchical optimization framework spanning model-level parallelism, computation-communication orchestration, and low-level kernel execution. The resulting system achieves 34.22% Model FLOPs Utilization (MFU) with a 2.93x improvement over the open-sou
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית