יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

ITC-MoE: חידוש בהתקן קיצור פרמטרים למודלי שפה MoE

ITC-MoE: Importance-guided Token-aware Compression for MoE Diffusion Language Models
מודלי שפה MoE: חידוש בהתקן קיצור פרמטרים. פיתוח חדש של LangGraph, Gemini ו-GPT-5.
תקציר מקורי באנגליתarXiv:2610.01296v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) Diffusion Language Models (DLMs) offer flexible parallel decoding and increased model capacity, but their large number of expert parameters incurs substantial computation and storage costs. Existing low-rank MoE compression methods largely rely on static factorization and fixed rank allocation, which overlook the distinctive properties of MoE DLMs. Specifically, we identify two properties: cross-mode non-uniform redundancy, where parameter redundancy and sensitivity to rank truncation vary across the input, output, and expert modes, and token-wise utilization variation, where hot and cold tokens exhibit distinct spectral characteristics and expert activation patterns. To address these challenges, we propose ITC-MoE, a
קרא במקור המקורי