כתבה
arXiv cs.LG ·
FedSubMuon: שיטה חדשה לקיצור זמן אימון מודלים
FedSubMuon: Communication-Efficient Federated LLM Fine-Tuning via Structured Subspace Muon
FedSubMuon היא שיטה חדשה לאימון מודלים בתפוצה, המשתמשת באופטימיזציה מתקדמת כדי לקצר זמן האימון. השיטה נבדקה על מודלים כמו Llama ו-Qwen, והראתה תוצאות טובות.
תקציר מקורי באנגליתarXiv:2609.06073v1 Announce Type: new Abstract: Federated fine-tuning adapts large language models (LLMs) to decentralized client data, but its scalability in cross-device training is often limited by the high communication cost. Muon is an optimizer that improves optimization performance by orthogonalizing momentum for matrix-valued parameters. Existing federated Muon methods demonstrate the benefit of matrix-aware optimization in federated learning, but still require transmitting full layer-size updates and optimizer state. A natural way to reduce communication is to directly apply Muon to LoRA factors, but this changes the optimized object and weakens Muon's matrix-aware update geometry. We propose FedSubMuon, a communication-efficient federated Muon fine-tuning method that optimizes co
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית