יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

MoARa: שיפור יעילות אימון מודלים LLM

MoARa: Module-Aware Rank Allocation and Structure-Preserving Decomposition for Low-Rank LLM Pre-training
MoARa משפרת יעילות אימון מודלים LLM גדולים. היא משלבת אלוקציה דינמית של דרגת פרויקציה ופירוק מבני. MoARa נבדקה על מודלים כמו Llama, Qwen ו-DeepSeek.
תקציר מקורי באנגליתarXiv:2609.15037v1 Announce Type: cross Abstract: Low-rank gradient projection reduces the optimizer-state memory cost of large language model (LLM) pretraining, but the steps and wall-clock time needed to reach a target quality remain a meaningful axis for improvement. We attribute this to two design choices in existing methods: the projection-rank budget is allocated uniformly across Transformer modules with heterogeneous projection sensitivity, and projecting a raw gradient attenuates its magnitude and direction jointly. We propose MoARa, which combines a static profiling-based module-aware projection-rank allocation with a block-wise magnitude-direction decomposition; the default block size is set in the neighborhood of the attention head dimension. Across five Transformer architecture
קרא במקור המקורי