כתבה
arXiv cs.AI ·
גזירה מובנית רב-מטרתית של LLM
Multi-Objective Structured Pruning of LLMs for Latency and Model Size Optimization
חוקרים הציגו שיטה חדשה לגזירה מובנית רב-מטרתית ל-Large Language Models. השיטה מטרה להפחית את גודל המודל ואת העיכוב, תוך שמירה על ביצועים טובים. הניסויים הראו תוצאות מבטיחות.
תקציר מקורי באנגליתarXiv:2607.22583v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved widespread adoption because of their strong reasoning and query-response capabilities. However, deploying them in embedded and edge computing environments remains challenging because of strict latency, memory, and energy constraints. Their large parameter counts and computational demands hinder efficient execution on resource-constrained platforms. Although model pruning has emerged as a viable solution for reducing scale while preserving performance, jointly optimizing layers, attention heads, and Multi-Layer Perceptron (MLP) dimensions remains highly complex. Exhaustively exploring this combined design space is computationally expensive and often leads to local optima or unstable configurations. To
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית