יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

UniRank: אלוקציה מאוחדת לדחיסת LLM

UniRank: Unified Rank Allocation for Low-Rank LLM Compression
UniRank היא שיטה לאלוקציה מאוחדת לדחיסת מודלי שפה גדולים. היא משלבת אנרגיה סינגולרית מקומית עם חשיבות פונקציונלית גלובלית, ומשפרת את דיוק הגיזרון ואת היכולת להידבקות במודלים גדולים.
תקציר מקורי באנגליתarXiv:2606.21847v2 Announce Type: replace-cross Abstract: Low-rank decomposition is a promising compression paradigm for large language models (LLMs), yet its effectiveness hinges on rank budget allocation across weight matrices: uniform or hand-crafted rules ignore module-wise importance, while learning-based allocation incurs substantial training overhead. We formulate rank allocation as a global sorting-and-truncation pipeline that scores every singular component by combining local singular energy with global functional importance, estimated via layer-wise input--output cosine similarity on a tiny calibration set. We show, both geometrically and empirically, that high input--output cosine similarity implies low effective rank. We further propose rank-preserving fine-tuning (RPFT), which
קרא במקור המקורי