כתבה
arXiv cs.AI ·
Train Overcomplete, Deploy Compact: Scaling Recovery Capacity for Structured LLM Pruning
תקציר מקורי באנגליתarXiv:2609.06974v1 Announce Type: cross Abstract: Large language models achieve strong performance across diverse tasks, but deployment remains costly because of memory, latency, and energy demands. Structured pruning reduces these costs by removing architectural components, yet its recovery stage is often limited by a mismatch between the recovery module's representational capacity and the complexity of the removed knowledge. We call this bottleneck the capacity-knowledge asymmetry and propose OverRep, an Overcomplete Reparameterization framework for structured LLM pruning. Following the principle of "train overcomplete, deploy compact", OverRep temporarily overparameterizes the recovery module during training to absorb complex knowledge distilled from the original model. After recovery,
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית