כתבה
arXiv cs.AI ·
Scaling Audio Models Efficiently: A Joint Study of Compute Constraints and Optimization Behavior
תקציר מקורי באנגליתarXiv:2606.22790v3 Announce Type: replace-cross Abstract: Large automatic speech recognition (ASR) models such as Whisper must be deployed across hardware with widely varying memory and inference-speed constraints. We present a compression framework that jointly parametrizes Whisper deployment along \emph{six} dimensions: model size $x_N$, temporal resolution $x_T$, encoder token stride $x_V$, low-rank adaptation capacity $x_R$, weight precision $x_Q$ and sparsity pattern $x_P$. All axes are jointly optimized against three deployment objectives (word error rate, inference FLOPs, and memory footprint) using a non-dominated sorting genetic evolutionary search (NSGA). Across 50 of the 1,680 candidate configurations evaluated, we measure the marginal effect of each axis on the three objectives
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית